Staff Observability Platform Engineer (AI / GPU Infrastructure)

New York, NY, US • Posted 10 hours ago • Updated 10 hours ago
Contract W2
12 Months
No Travel Required
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Kubernetes
  • Production Engineering
  • Staff Observability Platform Engineer
  • telemetry
  • Thanos
  • Elasticsearch
  • Prometheus
  • cardinality
  • Go
  • Python
  • AI/ML

Summary

Job Title: Staff Observability Platform Engineer (AI / GPU Infrastructure)

 Location: Seattle, WA / Houston, TX / New York, NY (Onsite)
 Employment Type: Contract


Position Overview

We are looking for a highly experienced Staff Observability Platform Engineer to design, build, and operate large-scale observability platforms supporting AI/GPU infrastructure and Kubernetes environments.

This is not a traditional monitoring or dashboard role. You''ll own the observability backend platform, ensuring it remains scalable, reliable, and cost-efficient as telemetry volumes grow.


Key Responsibilities

  • Design, build, and scale enterprise metricsloggingtracing, and telemetry platforms.
  • Architect and operate distributed observability backends using PrometheusMimirThanosVictoriaMetricsCortexLokiElasticsearch, or similar technologies.
  • Build and optimize OpenTelemetry Collector pipelines including routing, filtering, sampling, and exporters.
  • Optimize cardinalityingestionretentionstoragequery performance, and infrastructure cost.
  • Manage large-scale Kubernetes observability across multi-cluster environments.
  • Troubleshoot production issues involving metricslogstraces, networking, storage, and distributed systems.
  • Write and review production-quality code and establish observability standards across engineering teams.
  • Partner closely with Platform, Infrastructure, Security, and Application Engineering teams.

Required Qualifications

  • 8+ years of experience in ObservabilityPlatform EngineeringSRE, or Infrastructure Engineering.
  • Hands-on experience operating production-scale observability platforms such as MimirThanosVictoriaMetricsCortexLoki, or Elasticsearch.
  • Strong expertise with PrometheusOpenTelemetry, and telemetry pipeline design.
  • Experience managing large-scale production environments with measurable metrics (ingestion rates, active time series, storage, retention, cluster size, etc.).
  • Strong understanding of cardinalityretention strategiesstorage architecturesamplingquery optimization, and cost management.
  • Deep hands-on experience with Kubernetes, distributed systems, networking, service discovery, autoscaling, and reliability engineering.
  • Strong programming skills in Go and/orn Python.
  • Solid understanding of production engineering concepts including retries, backpressure, buffering, circuit breaking, graceful degradation, scalability, and failure handling.

Preferred Qualifications

  • Experience with GPU infrastructureAI/ML platforms, or HPC environments.
  • Knowledge of NVIDIA DCGMInfiniBandRoCE/RDMANVLinkNVSwitchNCCL, or Slurm.
  • Experience building custom Prometheus ExportersOpenTelemetry Collectors, or large multi-cluster Kubernetes observability platforms.

Ideal Candidate

We''re looking for someone who has personally owned and operated observability backends at production scale, understands the trade-offs between cardinality, retention, storage, performance, and cost, and can build scalable observability platforms for modern AI/GPU infrastructure. Experience limited to dashboards, alerts, or consuming monitoring tools without backend ownership will not be sufficient for this role.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91166803
  • Position Id: 9062384
  • Posted 10 hours ago
Contact the job poster
Vishal Puri

Vishal Puri

Sr Technical IT Recruiter @ Recruitment.ai
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

New York, New York

Today

Full-time

USD 275,000.00 - 377,000.00 per year

Jersey City, New Jersey

Today

Full-time

USD 149,000.00 - 186,000.00 per year

New York, New York

Today

Full-time

USD 149,000.00 - 186,000.00 per year

New York, New York

5d ago

Full-time

USD 149,000.00 - 186,000.00 per year

Search all similar jobs