Full time, on-site opportunity. Selected candidate will have to obtain and maintain an active US Top Secret Clearance.
The Senior AI Infrastructure Engineer will own the end-to-end stability, scalability, and resilience of our client's GPU training infrastructure. This role leads how the company trains at massive scale—designing, operating, and automating high-performance GPU clusters so ML research and platform teams can run reliably without manual intervention.
You’ll architect and maintain H200/B200/B300 and NVL72 systems, build and tune NVLink/InfiniBand/RoCE/Spectrum X fabrics, and integrate parallel storage platforms like VAST, DDN, and Weka to support terabyte scale multimodal workloads. A major focus is automated resilience: replacing manual triage with infrastructure as code, self-healing mechanisms, deep observability, and optimized scheduling across Kubernetes, Run:AI, and Ray.
The role blends hands-on datacenter work (rack/stack/cable/bring up) with high level platform ownership, including fleet health monitoring, fault isolation, congestion tuning, and onboarding engineers whose workloads depend on the cluster’s performance. You’ll partner across product and research teams to translate emerging compute needs into scalable platform capabilities.
REQUIREMENTS:
- 10+ years in a hands-on infrastructure, HPC, or datacenter engineering role supporting GPU compute at scale.
- Hands-on experience with H200/B200/B300 (or comparable) GPU systems: bring up, cabling, firmware/driver management.
- Experience with high-performance interconnects (NVLink, InfiniBand, RoCE, Spectrum-X) in clusters of hundreds of GPUs.
- Experience with high-performance parallel storage (VAST, DDN, Weka, Lustre, or similar).
- Kubernetes required; Run:ai or similar GPU scheduling/orchestration experience strongly preferred.
- Strong automation background. You build repeatable, automated deployment pipelines rather than manual processes.
- Able to lift/move 50+ lbs and perform physical datacenter work (rack/stack/cable/troubleshoot).
- Eligible to obtain and maintain an active U.S. Top Secret clearance.
PREFERRED QUALIFICATIONS
- Experience with NVIDIA NVL72 rack scale systems.
- Experience supporting LLM token serving/inference infrastructure alongside training clusters.
- Network fabric tuning experience (congestion control, adaptive routing, QoS) for RoCE/InfiniBand at scale.
- Familiarity with GPU/network observability tooling (DCGM, fabric telemetry) and automated fault detection.
- Experience supporting infrastructure as a shared platform serving multiple internal customer teams with differing requirements.