Senior AI Infrastructure Engineer

Costa Mesa, CA, US • Posted 1 day ago • Updated 1 day ago
Full Time
No Travel Required
On-site
$166,000 - $220,000/yr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • 10 yrs experience hands-on infrastructure HPC or datacenter supporting GPU compute
  • Hands-on with GPU systems including bringing uup and cabling aand firmwaare/driver mngt
  • High performance interconnect exp (NVLink
  • InfiniBand
  • RoCE or Spectrum-X
  • Exp with VAST or similar
  • Kubernetes
  • Automamtion
  • Ability to lift/move 50+ lbs

Summary

Full time, on-site opportunity.  Selected candidate will have to obtain and maintain an active US Top Secret Clearance. 

The Senior AI Infrastructure Engineer will own the end-to-end stability, scalability, and resilience of our client's GPU training infrastructure. This role leads how the company trains at massive scale—designing, operating, and automating high-performance GPU clusters so ML research and platform teams can run reliably without manual intervention.  

You’ll architect and maintain H200/B200/B300 and NVL72 systems, build and tune NVLink/InfiniBand/RoCE/Spectrum X fabrics, and integrate parallel storage platforms like VAST, DDN, and Weka to support terabyte scale multimodal workloads. A major focus is automated resilience: replacing manual triage with infrastructure as code, self-healing mechanisms, deep observability, and optimized scheduling across Kubernetes, Run:AI, and Ray.
 
The role blends hands-on datacenter work (rack/stack/cable/bring up) with high level platform ownership, including fleet health monitoring, fault isolation, congestion tuning, and onboarding engineers whose workloads depend on the cluster’s performance. You’ll partner across product and research teams to translate emerging compute needs into scalable platform capabilities.

REQUIREMENTS:

  • 10+ years in a hands-on infrastructure, HPC, or datacenter engineering role supporting GPU compute at scale.
  • Hands-on experience with H200/B200/B300 (or comparable) GPU systems: bring up, cabling, firmware/driver management.
  • Experience with high-performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of hundreds of GPUs.
  • Experience with high-performance parallel storage (VAST, DDN, Weka, Lustre, or similar).
  • Kubernetes required; Run:ai or similar GPU scheduling/orchestration experience strongly preferred.
  • Strong automation background. You build repeatable, automated deployment pipelines rather than manual processes.
  • Able to lift/move 50+ lbs and perform physical datacenter work (rack/stack/cable/troubleshoot).
  • Eligible to obtain and maintain an active U.S. Top Secret clearance.

PREFERRED QUALIFICATIONS

  • Experience with NVIDIA NVL72 rack scale systems.
  • Experience supporting LLM token serving/inference infrastructure alongside training clusters.
  • Network fabric tuning experience (congestion control, adaptive routing, QoS) for RoCE/InfiniBand at scale.
  • Familiarity with GPU/network observability tooling (DCGM, fabric telemetry) and automated fault detection.
  • Experience supporting infrastructure as a shared platform serving multiple internal customer teams with differing requirements.

 

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10238328
  • Position Id: 9105022
  • Posted 1 day ago
Contact the job poster
FC

Faith Chaskes

Recruiter @ Stone Search
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Costa Mesa, California

•

Today

Easy Apply

Full-time

$165,000 - $220,000

Hawthorne, California

•

Today

Full-time

USD 125,000.00 - 160,000.00 per year

Hawthorne, California

•

Today

Full-time

USD 165,000.00 - 265,000.00 per year

Hawthorne, California

•

Today

Full-time

USD 125,000.00 - 160,000.00 per year

Search all similar jobs