Senior AI Validation

Santa Clara, CA, US • Posted 21 hours ago • Updated 49 minutes ago
Contract W2
12 Months
No Travel Required
On-site
$60/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Artificial Intelligence
  • Communication
  • Firmware
  • Compatibility Testing
  • GPU
  • HPC
  • Ethernet

Summary

ENGAGEMENT SUMMARY

The Candidate will provide AI infrastructure validation services focused on proving cluster readiness, identifying

failure modes early, and accelerating root cause isolation before production impact. This role bridges systems,

networks, and workload behavior, and is ideal for a senior engineer who treats validation as an engineering

discipline rather than a checklist.

WHAT THIS CANDIDATE WILL BE DOING

• Design and execute validation plans for AI infrastructure spanning compute nodes, GPU communication,

fabric health, storage access, orchestration, and workload readiness.

• Run structured bring-up, soak, regression, and qualification tests on new or changed AI cluster

environments.

• Reproduce and isolate failures involving distributed training, node instability, communication libraries,

container stacks, storage paths, or network transport behavior.

• Build validation coverage for Ethernet and InfiniBand environments, including host readiness and end-to-

end workload verification.

• Correlate test failures with system logs, telemetry, firmware state, and application symptoms to accelerate

defect isolation.

• Partner with deployment, Linux, network, and platform teams to close validation gaps before operational

handoff.

• Create defect signatures, pass-fail criteria, readiness reports, and release recommendations.

• Improve automation for cluster certification, health scoring, and post-change validation.

WHAT WE NEE D TO SEE

• 7+ years in systems validation, performance engineering, QA for infrastructure, or AI/HPC environment

certification.

• Strong troubleshooting ability across Linux hosts, GPU systems, network fabrics, containers, and distributed

workload behavior.

• Experience designing validation strategies rather than only executing scripted test cases.

• Familiarity with AI workload dependencies such as NCCL, RDMA paths, storage throughput, and multi-node

orchestration behavior.

• Ability to distinguish infrastructure defects from workload, framework, or configuration issues.

• Strong scripting and automation capability for test execution and evidence collection.

• Clear written communication for readiness assessments and defect reports.

PREFERRED EXPERIENCE

• Experience validating GPU clusters, large training environments, or pre-production AI factories.

• Familiarity with telemetry analysis, burn-in workflows, and hardware-firmware-software compatibility

testing.

• Experience building qualification suites for both deployment gates and steady-state operations.

 

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10112463
  • Position Id: 9073095
  • Posted 21 hours ago
Contact the job poster
AJ

Amit Jha

Recruiter @ eTeam, Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Santa Clara, California

Yesterday

Easy Apply

Third Party, Contract

$57 - $70

Sunnyvale, California

Today

Full-time

USD 240,000.00 - 290,000.00 per year

Santa Clara, California

Today

Full-time

USD 70.00 - 75.00 per hour

San Jose, California

Today

Full-time

USD 210,000.00 per year

Search all similar jobs