Role :- Senior InfiniBand AI Network
Location :- Santa Clara, CA (Onsite)
Job Type : Contract
Job Description :
The Candidate will provide senior InfiniBand engineering services for AI and HPC clusters where fabric stability and latency-sensitive performance are mission critical. This is a hands-on role focused on cluster-scale bring-up, health validation, and deep troubleshooting of transport, fabric, and endpoint behavior.
Responsibilities :
Deploy and validate InfiniBand fabrics supporting distributed AI training and HPC workloads.
Troubleshoot issues involving fabric discovery, subnet management, link state, routing, partitioning, congestion, credit starvation, error counters, and host channel adapter behavior.
Diagnose job failures and performance degradation related to collective communication, NCCL transport selection, RDMA pathing, and fabric imbalance.
Validate switch, HCA, firmware, and cable consistency during cluster bring-up and expansion.
Use low-level fabric tooling to isolate bad links, flapping ports, unhealthy endpoints, topology mismatches, or subnet manager instability.
Partner with Linux, deployment, and AI validation teams to drive root cause analysis from application symptom to fabric source.
Define and execute pre-flight and post-change validation workflows for IB cluster readiness.
Document recurring fault patterns and create remediation playbooks for operational teams.
Required Skills :
7+ years in HPC or high-performance network environments, including direct InfiniBand operations experience.
Strong working knowledge of IB architecture, subnet management, link training, routing, partitions, congestion behavior, and performance diagnostics.
Experience with RDMA, NCCL-related network dependencies, and multi-node AI workload sensitivity to transport issues.
Strong troubleshooting skill using fabric health, counter, topology, and endpoint tools.
Experience with firmware and driver alignment across HCAs, switches, and Linux hosts.
Ability to triage complex issues that span host configuration, fabric state, and application communication behavior.
Preferred Skills :
Experience supporting DGX, GPU superpod, or equivalent AI cluster environments.
Familiarity with UFM, telemetry pipelines, and automated fabric validation.
Experience correlating IB anomalies with AI training performance outcomes.