Senior NVIDIA AI Factory / GPU Infrastructure Engineer NVIDIA H200

Remote • Posted 1 hour ago • Updated 1 hour ago
Full Time
No Travel Required
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Linux Administration
  • GPU
  • GPU cluster deployment
  • cuda
  • slurm
  • GPUDirect RDMA
  • NVIDIA
  • InfiniBand/NDR
  • RDMA
  • docker
  • kubernetes

Summary

We are looking for a Senior NVIDIA AI Factory / GPU Infrastructure Engineer to design, deploy, configure, and productionize a large-scale NVIDIA H200 GPU environment for an enterprise AI Factory initiative.
The engineer will work on an environment consisting of 32 NVIDIA H200 GPUs across HPE Gen12 GPU servers, high-speed InfiniBand networking, NVIDIA AI Enterprise, GPU workload orchestration, multi-tenant inference, NVIDIA NIM, and token/API-based AI services.
This is a hands-on role requiring strong experience taking GPU infrastructure from hardware readiness through production AI workloads. Physical rack/stack and power activities will be handled by the on-site infrastructure partner.
Key Responsibilities
GPU Infrastructure & Cluster Deployment
- Validate readiness of HPE DL380a Gen12 GPU servers with NVIDIA,  H200 GPUs.
- Perform GPU, NVLink, PCIe, firmware, BIOS, BMC and hardware health validation.
- Configure and maintain NVIDIA GPU drivers, CUDA, container runtime and supporting libraries.
- Validate GPU topology and multi-GPU communication.
- Configure and troubleshoot high-speed InfiniBand/NDR networking.
- Configure and validate GPUDirect RDMA and GPU-to-GPU communication.
- Perform NCCL and GPU communication/performance testing.
- Integrate high-performance storage for AI models, datasets and inference workloads.
- Establish validated firmware, OS, driver and CUDA baselines.
SLURM, MIG & GPU Resource Management
- Deploy and configure SLURM for GPU workload scheduling and orchestration.
- Configure NVIDIA Multi-Instance GPU (MIG) profiles where appropriate.
- Implement GPU resource allocation, scheduling, quotas and workload isolation.
- Support multi-user and multi-tenant GPU environments.
- Monitor GPU utilization, memory, workload performance and cluster health.
- Troubleshoot GPU scheduling, CUDA, NCCL, networking and performance issues.
AI Factory Architecture
- Build two isolated, highly available GPU clusters: one internal and one public-facing.
- Design production-grade GPU cluster architecture with security and workload isolation.
- Implement high availability, monitoring, logging and operational controls.
- Support capacity planning and GPU sizing for AI/LLM workloads.
- Develop deployment runbooks, architecture documentation and operational procedures.
NVIDIA AI Enterprise & NIM
- Deploy and configure NVIDIA AI Enterprise (NVAIE) components.
- Deploy LLMs and AI models using NVIDIA NIM.
- Configure scalable model-serving infrastructure.
- Optimize inference workloads for NVIDIA H200 GPUs.
- Perform model-serving benchmarking, capacity testing and performance tuning.
- Integrate inference endpoints with enterprise API gateways.
Multi-Tenant Inference Platform
- Design and implement a secure Inference-as-a-Service platform.
- Configure API-based access to hosted AI/LLM models.
- Implement tenant, application and model-level isolation.
- Support API key management and secure service access.
- Implement quotas, rate limits and GPU resource controls.
- Enable token-based usage metering and consumption tracking.
- Support model registry, prompt logging and AI governance controls.
Token Metering & API Gateway
- Integrate inference services with an enterprise API gateway.
- Capture input/output token consumption for LLM requests.
- Implement token metering, quotas and rate limiting.
- Enable per-tenant usage reporting and chargeback.
- Integrate usage data with billing/payment platforms.
- Support prepaid and postpaid AI consumption models.
Production Readiness
- Conduct end-to-end infrastructure and inference testing.
- Validate GPU performance, networking, storage and workload scheduling.
- Perform HA/failover and resiliency testing.
- Support security hardening and production-readiness assessments.
- Establish monitoring, alerting and operational dashboards.
- Define production acceptance criteria and SLA measurements.
- Provide technical documentation and knowledge transfer to client teams.
Required Skills
- 8+ years of Linux, infrastructure, HPC, cloud platform or systems engineering experience.
- Strong hands-on experience with NVIDIA GPU infrastructure.
- Experience deploying NVIDIA H100/H200 or comparable enterprise GPU platforms.
- Strong knowledge of:
  - NVIDIA CUDA
  - NVIDIA GPU Drivers
  - NVIDIA Container Toolkit
  - NCCL
  - GPUDirect RDMA
  - NVIDIA AI Enterprise
- Hands-on experience with SLURM.
- Experience with NVIDIA MIG and GPU resource partitioning.
- Strong Linux administration and troubleshooting skills.
- Experience with InfiniBand/RDMA high-performance networking.
- Understanding of GPU topology, PCIe, NUMA, NVLink and multi-GPU communication.
- Experience deploying containerized AI/ML workloads.
- Experience with Docker and Kubernetes/container orchestration technologies.
- Strong understanding of HA, monitoring, logging and production infrastructure practices.
AI/LLM Platform Experience
Candidates should have hands-on experience with several of the following:
- NVIDIA NIM
- NVIDIA AI Enterprise
- LLM inference infrastructure
- vLLM, Triton Inference Server or similar serving technologies
- API gateways
- Multi-tenant AI platforms
- Token metering and usage tracking
- Model registries
- AI governance and guardrails
- PrometheGrafana or equivalent observability platforms
- Python/Bash automation
Preferred Experience
- HPE ProLiant DL380/DL360 or similar enterprise GPU server platforms.
- NVIDIA H100/H200 deployments.
- Large-scale enterprise AI Factory implementations.
- Production LLM serving and inference optimization.
- Kubernetes GPU Operator.
- Infrastructure automation using Ansible/Terraform.
- Enterprise API gateway platforms.
- FinOps/chargeback platforms.
- Payment or billing-system integrations.
- Telecom or other large regulated enterprise environments.
- Security, data-sovereignty and compliance requirements.
Ideal Candidate
The ideal candidate is not solely an AI/ML developer or a traditional Linux administrator. We are looking for someone who understands the complete GPU-to-API stack:
HPE GPU Hardware → Linux/Firmware → NVIDIA Drivers/CUDA → InfiniBand/RDMA → SLURM/MIG → GPU Clusters → NVIDIA NIM → LLM Inference → API Gateway → Multi-Tenancy → Token Metering → Production Operations
The candidate should be comfortable independently troubleshooting infrastructure from the GPU and network layer through the model-serving and API layer.
Engagement
- Role: Senior NVIDIA AI Factory / GPU Infrastructure Engineer
- Project: Enterprise AI Factory / GPU Platform
- Initial Phase: Approximately 3–4 months
- Overall Program: Potential long-term engagement
- Environment: 32 × NVIDIA H200 GPUs across 4 GPU servers
- Work Model: Primarily remote with coordination with an on-site infrastructure partner
- Focus: GPU cluster deployment, SLURM/MIG, NVIDIA AI Enterprise, NIM, multi-tenant inference and production readiness
 
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90915265
  • Position Id: 9060681
  • Posted 1 hour ago
Contact the job poster
SM

Shipra Maheshwari

Recruiter @ LiveMindz
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Easy Apply

Contract

Depends on Experience

Remote or New York, New York

Today

Full-time

USD 250,000.00 - 485,000.00 per year

Remote

Today

Easy Apply

Contract

$90 - $100

Remote or Santa Clara, California

Today

Easy Apply

Third Party, Contract

Search all similar jobs