Senior GPU Infrastructure Engineer / NVIDIA AI Infrastructure Engineer

Remote • Posted 6 hours ago • Updated 6 hours ago
Contract W2
Contract Independent
Remote
$50 - $55/hr
Fitment

Dice Job Match Score™

📊 Calculating match score...

Job Details

Skills

  • NVIDIA H200
  • NVIDIA GPU Infrastructure
  • CUDA
  • NVIDIA AI Enterprise
  • NVIDIA NIM
  • SLURM
  • MIG
  • NCCL
  • GPUDirect RDMA
  • InfiniBand
  • Linux
  • GPU Cluster
  • NVLink
  • PCIe
  • NUMA
  • Docker
  • Kubernetes
  • GPU Operator
  • LLM Inference
  • vLLM
  • Triton Inference Server
  • Python
  • Bash
  • Prometheus
  • Grafana
  • API Gateway
  • Multi-Tenant AI
  • GPU Resource Management

Summary

Senior NVIDIA AI Factory / GPU Infrastructure Engineer NVIDIA H200

Position Overview

We are looking for a Senior NVIDIA AI Factory / GPU Infrastructure Engineer to design, deploy, configure, and productionize a large-scale NVIDIA H200 GPU environment for an enterprise AI Factory initiative.

The engineer will work on an environment consisting of 32 NVIDIA H200 GPUs across HPE Gen12 GPU servers, high-speed InfiniBand networking, NVIDIA AI Enterprise, GPU workload orchestration, multi-tenant inference, NVIDIA NIM, and token/API-based AI services.

This is a hands-on role requiring strong experience taking GPU infrastructure from hardware readiness through production AI workloads. Physical rack/stack and power activities will be handled by the on-site infrastructure partner.

Key Responsibilities

GPU Infrastructure & Cluster Deployment

- Validate readiness of HPE DL380a Gen12 GPU servers with NVIDIA H200 GPUs.
- Perform GPU, NVLink, PCIe, firmware, BIOS, BMC and hardware health validation.
- Configure and maintain NVIDIA GPU drivers, CUDA, container runtime and supporting libraries.
- Validate GPU topology and multi-GPU communication.
- Configure and troubleshoot high-speed InfiniBand/NDR networking.
- Configure and validate GPUDirect RDMA and GPU-to-GPU communication.
- Perform NCCL and GPU communication/performance testing.
- Integrate high-performance storage for AI models, datasets and inference workloads.
- Establish validated firmware, OS, driver and CUDA baselines.

SLURM, MIG & GPU Resource Management

- Deploy and configure SLURM for GPU workload scheduling and orchestration.
- Configure NVIDIA Multi-Instance GPU (MIG) profiles where appropriate.
- Implement GPU resource allocation, scheduling, quotas and workload isolation.
- Support multi-user and multi-tenant GPU environments.
- Monitor GPU utilization, memory, workload performance and cluster health.
- Troubleshoot GPU scheduling, CUDA, NCCL, networking and performance issues.

AI Factory Architecture

- Build two isolated, highly available GPU clusters: one internal and one public-facing.
- Design production-grade GPU cluster architecture with security and workload isolation.
- Implement high availability, monitoring, logging and operational controls.
- Support capacity planning and GPU sizing for AI/LLM workloads.
- Develop deployment runbooks, architecture documentation and operational procedures.

NVIDIA AI Enterprise & NIM

- Deploy and configure NVIDIA AI Enterprise (NVAIE) components.
- Deploy LLMs and AI models using NVIDIA NIM.
- Configure scalable model-serving infrastructure.
- Optimize inference workloads for NVIDIA H200 GPUs.
- Perform model-serving benchmarking, capacity testing and performance tuning.
- Integrate inference endpoints with enterprise API gateways.

Multi-Tenant Inference Platform

- Design and implement a secure Inference-as-a-Service platform.
- Configure API-based access to hosted AI/LLM models.
- Implement tenant, application and model-level isolation.
- Support API key management and secure service access.
- Implement quotas, rate limits and GPU resource controls.
- Enable token-based usage metering and consumption tracking.
- Support model registry, prompt logging and AI governance controls.

Token Metering & API Gateway

- Integrate inference services with an enterprise API gateway.
- Capture input/output token consumption for LLM requests.
- Implement token metering, quotas and rate limiting.
- Enable per-tenant usage reporting and chargeback.
- Integrate usage data with billing/payment platforms.
- Support prepaid and postpaid AI consumption models.

Production Readiness

- Conduct end-to-end infrastructure and inference testing.
- Validate GPU performance, networking, storage and workload scheduling.
- Perform HA/failover and resiliency testing.
- Support security hardening and production-readiness assessments.
- Establish monitoring, alerting and operational dashboards.
- Define production acceptance criteria and SLA measurements.
- Provide technical documentation and knowledge transfer to client teams.

Required Skills

- 8+ years of Linux, infrastructure, HPC, cloud platform or systems engineering experience.
- Strong hands-on experience with NVIDIA GPU infrastructure.
- Experience deploying NVIDIA H100/H200 or comparable enterprise GPU platforms.
- Strong knowledge of:
- NVIDIA CUDA
- NVIDIA GPU Drivers
- NVIDIA Container Toolkit
- NCCL
- GPUDirect RDMA
- NVIDIA AI Enterprise
- Hands-on experience with SLURM.
- Experience with NVIDIA MIG and GPU resource partitioning.
- Strong Linux administration and troubleshooting skills.
- Experience with InfiniBand/RDMA high-performance networking.
- Understanding of GPU topology, PCIe, NUMA, NVLink and multi-GPU communication.
- Experience deploying containerized AI/ML workloads.
- Experience with Docker and Kubernetes/container orchestration technologies.
- Strong understanding of HA, monitoring, logging and production infrastructure practices.

AI/LLM Platform Experience

Candidates should have hands-on experience with several of the following:

- NVIDIA NIM
- NVIDIA AI Enterprise
- LLM inference infrastructure
- vLLM, Triton Inference Server or similar serving technologies
- API gateways
- Multi-tenant AI platforms
- Token metering and usage tracking
- Model registries
- AI governance and guardrails
- PrometheGrafana or equivalent observability platforms
- Python/Bash automation

Preferred Experience

- HPE ProLiant DL380/DL360 or similar enterprise GPU server platforms.
- NVIDIA H100/H200 deployments.
- Large-scale enterprise AI Factory implementations.
- Production LLM serving and inference optimization.
- Kubernetes GPU Operator.
- Infrastructure automation using Ansible/Terraform.
- Enterprise API gateway platforms.
- FinOps/chargeback platforms.
- Payment or billing-system integrations.
- Telecom or other large regulated enterprise environments.
- Security, data-sovereignty and compliance requirements.

Ideal Candidate

The ideal candidate is not solely an AI/ML developer or a traditional Linux administrator. We are looking for someone who understands the complete GPU-to-API stack:

HPE GPU Hardware Linux/Firmware NVIDIA Drivers/CUDA InfiniBand/RDMA SLURM/MIG GPU Clusters NVIDIA NIM LLM Inference API Gateway Multi-Tenancy Token Metering Production Operations

The candidate should be comfortable independently troubleshooting infrastructure from the GPU and network layer through the model-serving and API layer.

Engagement

- Role: Senior NVIDIA AI Factory / GPU Infrastructure Engineer
- Project: Enterprise AI Factory / GPU Platform
- Initial Phase: Approximately 3 4 months
- Overall Program: Potential long-term engagement
- Environment: 32 NVIDIA H200 GPUs across 4 GPU servers
- Work Model: Primarily remote with coordination with an on-site infrastructure partner
- Focus: GPU cluster deployment, SLURM/MIG, NVIDIA AI Enterprise, NIM, multi-tenant inference and production readiness

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90915265
  • Position Id: 9055201
  • Posted 6 hours ago
Contact the job poster
AS

Anusha Santhosh

Recruiter @ LiveMindz
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote or New York, New York

Today

Full-time

USD 250,000.00 - 485,000.00 per year

Remote

Today

Full-time

Remote or California

Today

Full-time

USD 175,000.00 - 220,000.00 per year

Remote or Austin, Texas

Today

Full-time

$175,000 - $200,000 annually

Search all similar jobs