
Prudent Technologies and Consulting
Santa Clara, California • Today
Easy Apply
Contract
Depends on Experience
1500 results (58 new)

Prudent Technologies and Consulting
Santa Clara, California • Today
Easy Apply
Contract
Depends on Experience













Marvell Semiconductor Inc.
Santa Clara, California • Today
Full-time
USD 127,630.00 - 191,200.00 per year













Role: Senior AI Validation Engineer
Location: Santa Clara, CA (Onsite)
Job Type : Contract -W2
Job Description :
The Candidate will provide AI infrastructure validation services focused on proving cluster readiness, identifying failure modes early, and accelerating root cause isolation before production impact. This role bridges systems, networks, and workload behaviour, and is ideal for a senior engineer who treats validation as an engineering discipline rather than a checklist.
Responsibilities :
• Design and execute validation plans for AI infrastructure spanning compute nodes, GPU communication, fabric health, storage access, orchestration, and workload readiness.
• Run structured bring-up, soak, regression, and qualification tests on new or changed AI cluster environments.
• Reproduce and isolate failures involving distributed training, node instability, communication libraries, container stacks, storage paths, or network transport behavior.
• Build validation coverage for Ethernet and InfiniBand environments, including host readiness and end-to- end workload verification.
• Correlate test failures with system logs, telemetry, firmware state, and application symptoms to accelerate defect isolation.
• Partner with deployment, Linux, network, and platform teams to close validation gaps before operational handoff.
• Create defect signatures, pass-fail criteria, readiness reports, and release recommendations.
• Improve automation for cluster certification, health scoring, and post-change validation.
Required Skills :
• 10+ years in systems validation, performance engineering, QA for infrastructure, or AI/HPC environment certification.
• Strong troubleshooting ability across Linux hosts, GPU systems, network fabrics, containers, and distributed workload behavior.
• Experience designing validation strategies rather than only executing scripted test cases.
• Familiarity with AI workload dependencies such as NCCL, RDMA paths, storage throughput, and multi-node orchestration behavior.
• Ability to distinguish infrastructure defects from workload, framework, or configuration issues.
• Strong scripting and automation capability for test execution and evidence collection.
• Clear written communication for readiness assessments and defect reports.
Preferred Skills :
• Experience validating GPU clusters, large training environments, or pre-production AI factories.
• Familiarity with telemetry analysis, burn-in workflows, and hardware-firmware-software compatibility testing.
• Experience building qualification suites for both deployment gates and steady-state operations.
🔢 Crunching numbers...
Santa Clara, California
•
Today
Hi, Role :- Senior HPC Deployment and Cable Validation Location :- Santa Clara, CA(Onsite) P O S I T I O N D E S C R I P T I O N ENGAGEMENT SUMMARY The Candidate will provide deployment services for HPC and AI cluster infrastructure with an emphasis on physical build quality, cable validation, rack integration, and cluster readiness. This role is for a senior hands-on operator who can turn physical deployment work into repeatable, auditable infrastructure quality. WHAT THIS CANDIDATE WILL BE DOI
Easy Apply
Contract
$50 - $60
Santa Clara, California
•
Today
Role :- Senior InfiniBand AI Network Location :- Santa Clara, CA (Onsite) Job Type : Contract Job Description : The Candidate will provide senior InfiniBand engineering services for AI and HPC clusters where fabric stability and latency-sensitive performance are mission critical. This is a hands-on role focused on cluster-scale bring-up, health validation, and deep troubleshooting of transport, fabric, and endpoint behavior. Responsibilities : Deploy andvalidateInfiniBand fabrics supporting dist
Easy Apply
Contract, Third Party
Depends on Experience
Sunnyvale, California
•
Today
At Sonatus, we're driving the transformation to AI-enabled software-defined vehicles. Traditional automotive software methods can't keep pace with consumer expectations shaped by the mobile industry-where features evolve rapidly, update seamlessly, and improve continuously. That's why leading OEMs trust Sonatus to accelerate this shift. Our technology is already in production across more than 8 million vehicles on the road today and rapidly expanding. Headquartered in Sunnyvale, CA, with 250+ e
Full-time
USD 240,000.00 - 290,000.00 per year
Santa Clara, California
•
Yesterday
AI Infrastructure Engineer L3 Location: Santa Clara, CA Experience: 10 plus Years Employment Type: Full-Time About HCLTech HCLTech is a global technology company with over 220,000 professionals across 60 countries, delivering industry-leading capabilities in Digital, Engineering, Cloud, and AI. We help enterprises accelerate innovation through cutting-edge technologies and world-class talent. Job Summary We are seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, optimize, a
Easy Apply
Full-time
$120,000 - $160,000