Senior NVIDIA AI Factory / GPU Infrastructure Engineer NVIDIA H200
Position Overview
We are looking for a Senior NVIDIA AI Factory / GPU Infrastructure Engineer to design, deploy, configure, and productionize a large-scale NVIDIA H200 GPU environment for an enterprise AI Factory initiative.
The engineer will work on an environment consisting of 32 NVIDIA H200 GPUs across HPE Gen12 GPU servers, high-speed InfiniBand networking, NVIDIA AI Enterprise, GPU workload orchestration, multi-tenant inference, NVIDIA NIM, and token/API-based AI services.
This is a hands-on role requiring strong experience taking GPU infrastructure from hardware readiness through production AI workloads. Physical rack/stack and power activities will be handled by the on-site infrastructure partner.
Key Responsibilities
GPU Infrastructure & Cluster Deployment
- Validate readiness of HPE DL380a Gen12 GPU servers with NVIDIA H200 GPUs.
- Perform GPU, NVLink, PCIe, firmware, BIOS, BMC and hardware health validation.
- Configure and maintain NVIDIA GPU drivers, CUDA, container runtime and supporting libraries.
- Validate GPU topology and multi-GPU communication.
- Configure and troubleshoot high-speed InfiniBand/NDR networking.
- Configure and validate GPUDirect RDMA and GPU-to-GPU communication.
- Perform NCCL and GPU communication/performance testing.
- Integrate high-performance storage for AI models, datasets and inference workloads.
- Establish validated firmware, OS, driver and CUDA baselines.
SLURM, MIG & GPU Resource Management
- Deploy and configure SLURM for GPU workload scheduling and orchestration.
- Configure NVIDIA Multi-Instance GPU (MIG) profiles where appropriate.
- Implement GPU resource allocation, scheduling, quotas and workload isolation.
- Support multi-user and multi-tenant GPU environments.
- Monitor GPU utilization, memory, workload performance and cluster health.
- Troubleshoot GPU scheduling, CUDA, NCCL, networking and performance issues.
AI Factory Architecture
- Build two isolated, highly available GPU clusters: one internal and one public-facing.
- Design production-grade GPU cluster architecture with security and workload isolation.
- Implement high availability, monitoring, logging and operational controls.
- Support capacity planning and GPU sizing for AI/LLM workloads.
- Develop deployment runbooks, architecture documentation and operational procedures.
NVIDIA AI Enterprise & NIM
- Deploy and configure NVIDIA AI Enterprise (NVAIE) components.
- Deploy LLMs and AI models using NVIDIA NIM.
- Configure scalable model-serving infrastructure.
- Optimize inference workloads for NVIDIA H200 GPUs.
- Perform model-serving benchmarking, capacity testing and performance tuning.
- Integrate inference endpoints with enterprise API gateways.
Multi-Tenant Inference Platform
- Design and implement a secure Inference-as-a-Service platform.
- Configure API-based access to hosted AI/LLM models.
- Implement tenant, application and model-level isolation.
- Support API key management and secure service access.
- Implement quotas, rate limits and GPU resource controls.
- Enable token-based usage metering and consumption tracking.
- Support model registry, prompt logging and AI governance controls.
Token Metering & API Gateway
- Integrate inference services with an enterprise API gateway.
- Capture input/output token consumption for LLM requests.
- Implement token metering, quotas and rate limiting.
- Enable per-tenant usage reporting and chargeback.
- Integrate usage data with billing/payment platforms.
- Support prepaid and postpaid AI consumption models.
Production Readiness
- Conduct end-to-end infrastructure and inference testing.
- Validate GPU performance, networking, storage and workload scheduling.
- Perform HA/failover and resiliency testing.
- Support security hardening and production-readiness assessments.
- Establish monitoring, alerting and operational dashboards.
- Define production acceptance criteria and SLA measurements.
- Provide technical documentation and knowledge transfer to client teams.
Required Skills
- 8+ years of Linux, infrastructure, HPC, cloud platform or systems engineering experience.
- Strong hands-on experience with NVIDIA GPU infrastructure.
- Experience deploying NVIDIA H100/H200 or comparable enterprise GPU platforms.
- Strong knowledge of:
- NVIDIA CUDA
- NVIDIA GPU Drivers
- NVIDIA Container Toolkit
- NCCL
- GPUDirect RDMA
- NVIDIA AI Enterprise
- Hands-on experience with SLURM.
- Experience with NVIDIA MIG and GPU resource partitioning.
- Strong Linux administration and troubleshooting skills.
- Experience with InfiniBand/RDMA high-performance networking.
- Understanding of GPU topology, PCIe, NUMA, NVLink and multi-GPU communication.
- Experience deploying containerized AI/ML workloads.
- Experience with Docker and Kubernetes/container orchestration technologies.
- Strong understanding of HA, monitoring, logging and production infrastructure practices.
AI/LLM Platform Experience
Candidates should have hands-on experience with several of the following:
- NVIDIA NIM
- NVIDIA AI Enterprise
- LLM inference infrastructure
- vLLM, Triton Inference Server or similar serving technologies
- API gateways
- Multi-tenant AI platforms
- Token metering and usage tracking
- Model registries
- AI governance and guardrails
- PrometheGrafana or equivalent observability platforms
- Python/Bash automation
Preferred Experience
- HPE ProLiant DL380/DL360 or similar enterprise GPU server platforms.
- NVIDIA H100/H200 deployments.
- Large-scale enterprise AI Factory implementations.
- Production LLM serving and inference optimization.
- Kubernetes GPU Operator.
- Infrastructure automation using Ansible/Terraform.
- Enterprise API gateway platforms.
- FinOps/chargeback platforms.
- Payment or billing-system integrations.
- Telecom or other large regulated enterprise environments.
- Security, data-sovereignty and compliance requirements.
Ideal Candidate
The ideal candidate is not solely an AI/ML developer or a traditional Linux administrator. We are looking for someone who understands the complete GPU-to-API stack:
HPE GPU Hardware Linux/Firmware NVIDIA Drivers/CUDA InfiniBand/RDMA SLURM/MIG GPU Clusters NVIDIA NIM LLM Inference API Gateway Multi-Tenancy Token Metering Production Operations
The candidate should be comfortable independently troubleshooting infrastructure from the GPU and network layer through the model-serving and API layer.
Engagement
- Role: Senior NVIDIA AI Factory / GPU Infrastructure Engineer
- Project: Enterprise AI Factory / GPU Platform
- Initial Phase: Approximately 3 4 months
- Overall Program: Potential long-term engagement
- Environment: 32 NVIDIA H200 GPUs across 4 GPU servers
- Work Model: Primarily remote with coordination with an on-site infrastructure partner
- Focus: GPU cluster deployment, SLURM/MIG, NVIDIA AI Enterprise, NIM, multi-tenant inference and production readiness