Senior InfiniBand AI Network

Santa Clara, CA, US • Posted 1 day ago • Updated 5 hours ago
Contract W2
Contract Corp To Corp
Contract Independent
12 Months
No Travel Required
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Artificial Intelligence
  • HPC
  • GPU
  • Remote Direct Memory Access
  • InfiniBand

Summary

Role :- Senior InfiniBand AI Network
Location :- Santa Clara, CA (Onsite)
Job Type : Contract


Job Description :

The Candidate will provide senior InfiniBand engineering services for AI and HPC clusters where fabric stability and latency-sensitive performance are mission critical. This is a hands-on role focused on cluster-scale bring-up, health validation, and deep troubleshooting of transport, fabric, and endpoint behavior.

Responsibilities :

Deploy and validate InfiniBand fabrics supporting distributed AI training and HPC workloads.

Troubleshoot issues involving fabric discovery, subnet management, link state, routing, partitioning, congestion, credit starvation, error counters, and host channel adapter behavior.

Diagnose job failures and performance degradation related to collective communication, NCCL transport selection, RDMA pathing, and fabric imbalance.

Validate switch, HCA, firmware, and cable consistency during cluster bring-up and expansion.

Use low-level fabric tooling to isolate bad links, flapping ports, unhealthy endpoints, topology mismatches, or subnet manager instability.

Partner with Linux, deployment, and AI validation teams to drive root cause analysis from application symptom to fabric source.

Define and execute pre-flight and post-change validation workflows for IB cluster readiness.

Document recurring fault patterns and create remediation playbooks for operational teams.

 

Required Skills :

 

7+ years in HPC or high-performance network environments, including direct InfiniBand operations experience.

Strong working knowledge of IB architecture, subnet management, link training, routing, partitions, congestion behavior, and performance diagnostics.

Experience with RDMA, NCCL-related network dependencies, and multi-node AI workload sensitivity to transport issues.

Strong troubleshooting skill using fabric health, counter, topology, and endpoint tools.

Experience with firmware and driver alignment across HCAs, switches, and Linux hosts.

Ability to triage complex issues that span host configuration, fabric state, and application communication behavior.

Preferred Skills :

Experience supporting DGX, GPU superpod, or equivalent AI cluster environments.

Familiarity with UFM, telemetry pipelines, and automated fabric validation.

Experience correlating IB anomalies with AI training performance outcomes.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: prutx001
  • Position Id: 9068892
  • Posted 1 day ago
Contact the job poster
PK

Pavan Kalva

Recruiter @ Prudent Technologies and Consulting
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Santa Clara, California

Yesterday

Easy Apply

Third Party, Contract

$58 - $60

Santa Clara, California

Today

Easy Apply

Contract

Depends on Experience

Santa Clara, California

Today

Easy Apply

Contract

$50 - $60

Santa Clara, California

2d ago

Easy Apply

Contract

Depends on Experience

Search all similar jobs