AI Senior Engineer
Plano, TX
Fulltime permanent
About the role
Two halves, and both matter. You make the platform useful — enabling distributed training and production inference on the clusters we build. And you own our infrastructure validation offering: the independent, evidence-based assessment that tells a customer whether the eight-figure GPU estate they just bought actually performs the way the reference architecture promised.
If you like building the benchmark that settles the argument, this is your role.
What you’ll do
Workload engineering - Enable and tune distributed training: PyTorch DDP and FSDP, DeepSpeed/ZeRO, tensor and pipeline parallelism, NCCL tuning, MFU measurement and improvement. - Architect inference serving — Triton, NVIDIA NIM, vLLM, TensorRT-LLM — with quantization, continuous batching, KV-cache and paged attention, and autoscaling to latency SLOs. - Size inference platforms backward from customer SLOs: time-to-first-token, inter-token latency, concurrency, context length. - Support the GenAI patterns customers actually ask for — RAG, fine-tuning, agentic pipelines, vector database integration.
Validation offering - Build and own our GPU infrastructure validation suite: fabric validation, NCCL scaling curves, GPU burn and thermal/power soak, storage throughput, GPUDirect verification, reference-workload MFU baselines. - Deliver acceptance validation engagements and author the signed acceptance reports customers use to hold vendors to their commitments. - Design the continuous-validation service: post-downtime revalidation, driver and firmware matrix certification, performance drift detection against baseline. - Produce published benchmark methodology and reports that establish our technical credibility in the market. - Own PoC workload design and the benchmark evidence that closes engagements.
What you need
5+ years ML engineering, MLOps, or performance engineering with production GPU workloads.
Hands-on distributed training on multi-node GPU clusters — not single-GPU or notebook-scale work.
Production LLM inference serving: Triton, vLLM, or TensorRT-LLM.
Rigorous benchmarking discipline — you can profile a system, explain where GPU time actually goes, and defend a number.
Strong Python; containers and Kubernetes-native workflows (Kubeflow, Ray, Argo Workflows).
Clear technical writing. A validation report is a deliverable a customer pays for.
Nice to have
NVIDIA NIM and NeMo; MLPerf or formal benchmark program experience; test-and-validation engineering background; fine-tuning at scale (LoRA/QLoRA); customer-facing PoC delivery.