Senior Engineer AI Workload & Infrastructure Validation

Hybrid in Plano, TX, US • Posted 2 hours ago • Updated 2 hours ago
Full Time
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • ML Engineering
  • MLOps
  • GPU Workloads
  • Multi-Node GPU Clusters
  • Distributed Training
  • PyTorch
  • DDP
  • FSDP
  • DeepSpeed
  • ZeRO
  • Tensor Parallelism
  • Pipeline Parallelism
  • NCCL
  • MFU
  • LLM Inference
  • Triton
  • NVIDIA NIM
  • vLLM
  • TensorRT-LLM
  • Quantization
  • Continuous Batching
  • KV-Cache
  • Paged Attention
  • Autoscaling
  • RAG
  • Fine-Tuning
  • Agentic AI
  • Vector Databases
  • Python
  • Kubernetes
  • Containers
  • Kubeflow
  • Ray
  • Argo Workflows
  • GPU Benchmarking
  • GPU Validation
  • GPUDirect
  • Performance Engineering
  • Technical Writing

Summary

Job Title: Senior Engineer AI Workload & Infrastructure Validation
Client: LTTS
Location: Plano, TX Hybrid Employment Type: FTE
Experience: 5+ Years ML Engineering, MLOps, or Performance Engineering with Production GPU Workloads Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service

About the Role

We are building a GPU-as-a-Service and AI Factory practice supporting enterprise and industrial customers. This role combines AI workload engineering, distributed training, production inference, GPU infrastructure validation, benchmarking, and customer acceptance testing.

You will enable and tune distributed training and inference workloads while owning an independent validation offering that determines whether GPU infrastructure performs according to the reference architecture.

Key Responsibilities

  • Enable and tune distributed training using PyTorch DDP, FSDP, DeepSpeed/ZeRO, tensor parallelism, pipeline parallelism, and NCCL tuning.
  • Measure and improve Model FLOPs Utilization (MFU) across distributed GPU workloads.
  • Architect production inference serving using Triton, NVIDIA NIM, vLLM, and TensorRT-LLM.
  • Implement and optimize quantization, continuous batching, KV-cache, paged attention, and autoscaling to meet latency SLOs.
  • Size inference platforms based on customer time-to-first-token, inter-token latency, concurrency, and context-length requirements.
  • Support GenAI workloads including RAG, fine-tuning, agentic pipelines, and vector database integration.
  • Build and maintain a GPU infrastructure validation suite covering fabric validation, NCCL scaling, GPU burn/thermal/power testing, storage throughput, GPUDirect verification, and workload MFU baselines.
  • Deliver GPU infrastructure acceptance validation engagements and prepare customer-facing acceptance reports.
  • Design continuous validation processes covering driver/firmware certification and performance drift detection.
  • Develop benchmark methodologies, reports, and evidence for customer engagements.
  • Own PoC workload design and benchmark execution.

Required Qualifications

  • 5+ years of ML engineering, MLOps, or performance engineering experience with production GPU workloads.
  • Hands-on experience with distributed training on multi-node GPU clusters.
  • Production LLM inference serving experience with Triton, vLLM, or TensorRT-LLM.
  • Strong benchmarking and performance-analysis skills with the ability to profile systems and explain GPU utilization and performance results.
  • Strong Python development experience.
  • Experience with containers and Kubernetes-native workflows, including Kubeflow, Ray, or Argo Workflows.
  • Strong technical writing skills for producing customer-facing validation reports.

Nice to Have

  • NVIDIA NIM and NeMo experience.
  • MLPerf or formal benchmark program experience.
  • Test-and-validation engineering background.
  • Large-scale fine-tuning using LoRA/QLoRA.
  • Customer-facing PoC delivery experience.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10202400
  • Position Id: 9088309
  • Posted 2 hours ago
Contact the job poster
DG

Dolly Gupta

Recruiter @ VST Consulting, Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Full-time

Remote

Today

Full-time

USD 296,300.00 - 453,900.00 per year

No location provided

Today

Full-time

USD 79,200.00 per year

Remote or California

Today

Full-time

USD 218.00 - 400.00 per day

Search all similar jobs