Job Title: Staff Observability Platform Engineer (AI / GPU Infrastructure)
Location: Seattle, WA / Houston, TX / New York, NY (Onsite)
Employment Type: Contract
Position Overview
We are looking for a highly experienced Staff Observability Platform Engineer to design, build, and operate large-scale observability platforms supporting AI/GPU infrastructure and Kubernetes environments.
This is not a traditional monitoring or dashboard role. You''ll own the observability backend platform, ensuring it remains scalable, reliable, and cost-efficient as telemetry volumes grow.
Key Responsibilities
- Design, build, and scale enterprise metrics, logging, tracing, and telemetry platforms.
- Architect and operate distributed observability backends using Prometheus, Mimir, Thanos, VictoriaMetrics, Cortex, Loki, Elasticsearch, or similar technologies.
- Build and optimize OpenTelemetry Collector pipelines including routing, filtering, sampling, and exporters.
- Optimize cardinality, ingestion, retention, storage, query performance, and infrastructure cost.
- Manage large-scale Kubernetes observability across multi-cluster environments.
- Troubleshoot production issues involving metrics, logs, traces, networking, storage, and distributed systems.
- Write and review production-quality code and establish observability standards across engineering teams.
- Partner closely with Platform, Infrastructure, Security, and Application Engineering teams.
Required Qualifications
- 8+ years of experience in Observability, Platform Engineering, SRE, or Infrastructure Engineering.
- Hands-on experience operating production-scale observability platforms such as Mimir, Thanos, VictoriaMetrics, Cortex, Loki, or Elasticsearch.
- Strong expertise with Prometheus, OpenTelemetry, and telemetry pipeline design.
- Experience managing large-scale production environments with measurable metrics (ingestion rates, active time series, storage, retention, cluster size, etc.).
- Strong understanding of cardinality, retention strategies, storage architecture, sampling, query optimization, and cost management.
- Deep hands-on experience with Kubernetes, distributed systems, networking, service discovery, autoscaling, and reliability engineering.
- Strong programming skills in Go and/orn Python.
- Solid understanding of production engineering concepts including retries, backpressure, buffering, circuit breaking, graceful degradation, scalability, and failure handling.
Preferred Qualifications
- Experience with GPU infrastructure, AI/ML platforms, or HPC environments.
- Knowledge of NVIDIA DCGM, InfiniBand, RoCE/RDMA, NVLink, NVSwitch, NCCL, or Slurm.
- Experience building custom Prometheus Exporters, OpenTelemetry Collectors, or large multi-cluster Kubernetes observability platforms.
Ideal Candidate
We''re looking for someone who has personally owned and operated observability backends at production scale, understands the trade-offs between cardinality, retention, storage, performance, and cost, and can build scalable observability platforms for modern AI/GPU infrastructure. Experience limited to dashboards, alerts, or consuming monitoring tools without backend ownership will not be sufficient for this role.