Client is seeking a Sr AI Validation Engineer to provide artificial intelligence infrastructure validation services focused on proving cluster readiness, identifying failure modes early, and accelerating root cause isolation before production impact. This role bridges systems, networks, and workload behavior and is ideal for a senior engineer who approaches validation as an engineering discipline rather than a checklist driven process.
Responsibilities:
Design and execute validation plans for artificial intelligence infrastructure spanning compute nodes, GPU communication, fabric health, storage access, orchestration, and workload readiness.
Run structured bring up, soak, regression, and qualification testing within new or modified artificial intelligence cluster environments.
Reproduce and isolate failures involving distributed training, node instability, communication libraries, container platforms, storage paths, or network transport behavior.
Build validation coverage for Ethernet and InfiniBand environments, including host readiness and end to end workload verification.
Correlate test failures with system logs, telemetry, firmware state, and application symptoms to accelerate defect isolation.
Partner with deployment, Linux, networking, and platform teams to close validation gaps prior to operational handoff.
Create defect signatures, pass fail criteria, readiness reports, and release recommendations.
Improve automation for cluster certification, health scoring, and post change validation activities. Qualifications:
Required Qualifications
7 or more years of experience in systems validation, performance engineering, infrastructure quality assurance, or artificial intelligence and high performance computing environment certification.
Strong troubleshooting ability across Linux hosts, GPU systems, network fabrics, containers, and distributed workload behavior.
Experience designing validation strategies and frameworks rather than solely executing pre existing test cases.
Familiarity with artificial intelligence workload dependencies including NCCL, RDMA paths, storage throughput, and multi node orchestration behavior.
Ability to distinguish infrastructure defects from workload, framework, or configuration issues.
Strong scripting and automation capabilities for test execution, evidence collection, and reporting.
Excellent written communication skills with experience creating readiness assessments and detailed defect reports.
Preferred Qualifications
Experience validating GPU clusters, large scale training environments, or artificial intelligence infrastructure environments before production deployment.
Familiarity with telemetry analysis, burn in workflows, and hardware firmware software compatibility testing.
Experience building qualification suites for both deployment readiness gates and steady state operational health.
Tools and Technologies:
Linux
GPU Infrastructure
NCCL
RDMA
Kubernetes
Container Technologies
Ethernet Networks
InfiniBand Fabrics
Python
Bash
Automation Frameworks
Telemetry Platforms
Distributed Training Systems
Performance Analysis Tools
Infrastructure Validation Tools