Job Summary
We are seeking a highly experienced Senior AI Infrastructure Engineer to deploy, operate, secure, and optimize NVIDIA DGX-based AI infrastructure supporting large-scale AI/ML training and inference workloads.
The ideal candidate will have strong hands-on expertise across NVIDIA DGX systems, Kubernetes, NVIDIA GPU Operator, InfiniBand, BlueField DPUs, NVLink/NVSwitch, and AI infrastructure automation. This is a deeply technical role requiring experience managing high-performance GPU clusters and secure, scalable Kubernetes environments.
Key Responsibilities
AI Infrastructure Operations
- Deploy and manage NVIDIA DGX BasePOD and SuperPOD environments.
- Manage DGX lifecycle operations including provisioning, monitoring, firmware upgrades, and capacity planning.
- Use Base Command Manager for GPU cluster management and workload orchestration.
- Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification.
Kubernetes & AI Platform Engineering
- Design and manage secure, scalable Kubernetes clusters optimized for GPU workloads.
- Configure and manage NVIDIA GPU Operator.
- Deploy and manage AI/ML workloads using Kubernetes, Helm, and Kubeflow.
- Implement CI/CD and GitOps practices for ML workflows and infrastructure.
High-Performance Networking
- Administer InfiniBand fabrics and BlueField DPUs.
- Manage InfiniBand infrastructure using Unified Fabric Manager (UFM).
- Configure and optimize NVLink/NVSwitch connectivity and performance.
- Leverage BlueField DPUs for storage, firewalling, security, and telemetry offload.
Security & Compliance
- Apply CKS-level security practices to Kubernetes and containerized AI environments.
- Implement RBAC, workload identity, secrets management, network segmentation, and auditing.
- Support zero-trust security initiatives and container/model supply-chain security.
Monitoring & Optimization
- Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, and Grafana.
- Optimize GPU utilization, AI workload performance, and infrastructure efficiency.
- Develop operational runbooks, incident response procedures, and SLA dashboards.
Required Skills & Experience
- Strong hands-on experience with NVIDIA DGX, BasePOD, and SuperPOD environments.
- Experience with Kubernetes, NVIDIA GPU Operator, Helm, and Kubeflow.
- Proven experience administering InfiniBand and UFM.
- Hands-on experience with NVIDIA BlueField DPUs.
- Experience with Base Command Manager.
- Strong knowledge of NVLink/NVSwitch and NCCL.
- Strong scripting and automation skills using Python, YAML, and Bash.
- Experience securing GPU-enabled Kubernetes environments.
- CKA, CKAD, and CKS certifications.
- NVIDIA certifications such as NCA-AIIO, NCP-AII, NCP-AIO, and NCP-AIN.
Preferred Qualifications
- Experience with Ansible, Terraform, GitOps, and CI/CD.
- Experience with NFS, BeeGFS, or Lustre storage.
- Knowledge of RoCE, InfiniBand, RDMA, gRPC, and DPU offload.
- Experience supporting large-scale AI/ML infrastructure and MLOps environments.
Technical Environment
AI/GPU: NVIDIA DGX, BasePOD, SuperPOD, NVLink, NVSwitch, NCCL, NVIDIA DCGM
Kubernetes: Kubernetes, GPU Operator, Helm, Kubeflow
Networking: InfiniBand, UFM, BlueField DPU, RoCE, RDMA
Automation: Python, Bash, YAML, Terraform, Ansible
DevOps: GitOps, CI/CD
Monitoring: Prometheus, Grafana
Storage: NFS, BeeGFS, Lustre