Job Summary
We are seeking a Principal AI Infrastructure Architect to design, build, and operate secure, scalable GPU-accelerated AI platforms. The ideal candidate will have deep expertise in Kubernetes, NVIDIA DGX infrastructure, InfiniBand networking, BlueField DPUs, and MLOps platforms.
This is a hands-on architecture role responsible for building production-grade AI infrastructure that supports large-scale ML training and inference workloads.
Key Responsibilities
Kubernetes & AI Platform Architecture
- Architect and manage Kubernetes-based AI/ML platforms running on NVIDIA DGX systems.
- Integrate NVIDIA Base Command Manager with Kubernetes for GPU workload scheduling and resource optimization.
- Design MIG-based GPU partitioning strategies for multi-tenant environments.
- Develop and manage Helm charts, custom controllers, and GPU operators.
DGX Infrastructure & Capacity Planning
- Administer and optimize NVIDIA DGX BasePOD and SuperPOD environments.
- Ensure optimal GPU, CPU, storage, and cluster performance.
- Manage DGX system lifecycle, updates, and infrastructure operations.
- Lead capacity planning for cluster expansion, including power, cooling, and storage requirements.
High-Performance Networking
- Deploy and manage InfiniBand fabrics using Unified Fabric Manager (UFM).
- Configure and operate NVIDIA BlueField DPUs for networking, security, and storage offload.
- Optimize data pipelines between storage and GPU infrastructure.
Security & Compliance
- Apply CKS-level security practices to Kubernetes clusters and AI workloads.
- Implement service mesh, microsegmentation, and secure workload architectures.
- Conduct vulnerability assessments, audits, and policy enforcement.
Automation, Observability & MLOps
- Automate infrastructure provisioning using Terraform and Ansible.
- Develop automation using Python, Bash, and YAML.
- Monitor AI infrastructure using NVIDIA DCGM, Prometheus, and Grafana.
- Build MLOps workflows using Kubeflow Pipelines and NVIDIA Triton Inference Server.
- Troubleshoot complex issues across hardware, networking, Kubernetes, and AI software layers.
Required Skills & Experience
- Hands-on experience running AI/ML workloads on NVIDIA DGX systems.
- Strong expertise in Kubernetes administration, architecture, and security.
- Deep experience with InfiniBand, UFM, and BlueField DPU administration.
- Strong scripting and automation skills with Python, Bash, and YAML.
- Experience designing scalable, secure, production-grade AI infrastructure.
- CKA, CKAD, and CKS certifications.
Preferred Qualifications
- Experience with NVIDIA Base Command Manager.
- Experience with NVIDIA GPU Operator.
- Experience with Kubeflow/Kubeflow Pipelines.
- Knowledge of MIG GPU partitioning.
- Experience with Terraform and Ansible.
- Experience with NVIDIA Triton Inference Server.
- Strong collaboration skills with ML researchers, DevOps engineers, and infrastructure teams.
Technical Environment
GPU Infrastructure: NVIDIA DGX, BasePOD, SuperPOD, NVIDIA AI Enterprise
Kubernetes: Kubernetes, GPU Operator, Helm, Custom Controllers, MIG
Networking: InfiniBand, UFM, BlueField DPU
MLOps: Kubeflow Pipelines, NVIDIA Triton Inference Server
Automation: Terraform, Ansible, Python, Bash, YAML
Monitoring: NVIDIA DCGM, Prometheus, Grafana