Senior AI Infrastructure Engineer

Remote • Posted 5 hours ago • Updated 5 hours ago
Contract W2
12 Months
No Travel Required
Remote
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • NVIDIA DGX / BasePOD / SuperPOD
  • Kubernetes & NVIDIA GPU Operator
  • InfiniBand & BlueField DPU
  • NVLink / NVSwitch & NCCL

Summary

Job Summary

We are seeking a highly experienced Senior AI Infrastructure Engineer to deploy, operate, secure, and optimize NVIDIA DGX-based AI infrastructure supporting large-scale AI/ML training and inference workloads.

The ideal candidate will have strong hands-on expertise across NVIDIA DGX systems, Kubernetes, NVIDIA GPU Operator, InfiniBand, BlueField DPUs, NVLink/NVSwitch, and AI infrastructure automation. This is a deeply technical role requiring experience managing high-performance GPU clusters and secure, scalable Kubernetes environments.

Key Responsibilities

AI Infrastructure Operations

  • Deploy and manage NVIDIA DGX BasePOD and SuperPOD environments.
  • Manage DGX lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning.
  • Use Base Command Manager for GPU cluster management and workload orchestration.
  • Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification.

Kubernetes & AI Platform Engineering

  • Design and manage secure, scalable Kubernetes clusters optimized for GPU workloads.
  • Configure and manage NVIDIA GPU Operator.
  • Deploy and manage AI/ML workloads using Kubernetes, Helm, and Kubeflow.
  • Implement CI/CD and GitOps practices for ML workflows and infrastructure.

High-Performance Networking

  • Administer InfiniBand fabrics and BlueField DPUs.
  • Manage InfiniBand infrastructure using Unified Fabric Manager (UFM).
  • Configure and optimize NVLink/NVSwitch connectivity and performance.
  • Leverage BlueField DPUs for storage, firewalling, security, and telemetry offload.

Security & Compliance

  • Apply CKS-level security practices to Kubernetes and containerized AI environments.
  • Implement RBAC, workload identity, secrets management, network segmentation, and auditing.
  • Support zero-trust security initiatives and container/model supply-chain security.

Monitoring & Optimization

  • Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, and Grafana.
  • Optimize GPU utilization, AI workload performance, and infrastructure efficiency.
  • Develop operational runbooks, incident response procedures, and SLA dashboards.

Required Skills & Experience

  • Strong hands-on experience with NVIDIA DGX, BasePOD, and SuperPOD environments.
  • Experience with Kubernetes, NVIDIA GPU Operator, Helm, and Kubeflow.
  • Proven experience administering InfiniBand and UFM.
  • Hands-on experience with NVIDIA BlueField DPUs.
  • Experience with Base Command Manager.
  • Strong knowledge of NVLink/NVSwitch and NCCL.
  • Strong scripting and automation skills using Python, YAML, and Bash.
  • Experience securing GPU-enabled Kubernetes environments.
  • CKA, CKAD, and CKS certifications.
  • NVIDIA certifications such as NCA-AIIO, NCP-AII, NCP-AIO, and NCP-AIN.

Preferred Qualifications

  • Experience with Ansible, Terraform, GitOps, and CI/CD.
  • Experience with NFS, BeeGFS, or Lustre storage.
  • Knowledge of RoCE, InfiniBand, RDMA, gRPC, and DPU offload.
  • Experience supporting large-scale AI/ML infrastructure and MLOps environments.

Technical Environment

AI/GPU: NVIDIA DGX, BasePOD, SuperPOD, NVLink, NVSwitch, NCCL, NVIDIA DCGM
Kubernetes: Kubernetes, GPU Operator, Helm, Kubeflow
Networking: InfiniBand, UFM, BlueField DPU, RoCE, RDMA
Automation: Python, Bash, YAML, Terraform, Ansible
DevOps: GitOps, CI/CD
Monitoring: Prometheus, Grafana
Storage: NFS, BeeGFS, Lustre

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10371609
  • Position Id: 9064157
  • Posted 5 hours ago

Company Info

About Prabhav Services Inc

High ROI

Many companies find that constant maintenance eats into their budget for new technology. By outsourcing your IT management to us, you can focus on what you do best--running your business.

Satisfaction Guaranteed

That's why our goal is to provide an experience that is tailored to your company's needs. No matter the budget, we pride ourselves on providing professional customer service.

Technical Experience

We are well-versed in a variety of operating systems, networks, and databases. We use this expertise to help our customers with a variety of small to mid-sized projects.

Contact the job poster
AS

Aniket Sharma

Recruiter @ Prabhav Services Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Easy Apply

Contract

Depends on Experience

Remote

Today

Easy Apply

Contract

Depends on Experience

Remote

Today

Easy Apply

Contract

Depends on Experience

Search all similar jobs