Principal AI Infrastructure Architect

Remote • Posted 5 hours ago • Updated 5 hours ago
Contract W2
12 Months
No Travel Required
Remote
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • NVIDIA DGX / BasePOD / SuperPOD
  • Kubernetes Architecture & Security
  • InfiniBand / UFM / BlueField DPU
  • GPU Infrastructure & MLOps — GPU Operator

Summary

Job Summary

We are seeking a Principal AI Infrastructure Architect to design, build, and operate secure, scalable GPU-accelerated AI platforms. The ideal candidate will have deep expertise in Kubernetes, NVIDIA DGX infrastructure, InfiniBand networking, BlueField DPUs, and MLOps platforms.

This is a hands-on architecture role responsible for building production-grade AI infrastructure that supports large-scale ML training and inference workloads.

Key Responsibilities

Kubernetes & AI Platform Architecture

  • Architect and manage Kubernetes-based AI/ML platforms running on NVIDIA DGX systems.
  • Integrate NVIDIA Base Command Manager with Kubernetes for GPU workload scheduling and resource optimization.
  • Design MIG-based GPU partitioning strategies for multi-tenant environments.
  • Develop and manage Helm charts, custom controllers, and GPU operators.

DGX Infrastructure & Capacity Planning

  • Administer and optimize NVIDIA DGX BasePOD and SuperPOD environments.
  • Ensure optimal GPU, CPU, storage, and cluster performance.
  • Manage DGX system lifecycle, updates, and infrastructure operations.
  • Lead capacity planning for cluster expansion, including power, cooling, and storage requirements.

High-Performance Networking

  • Deploy and manage InfiniBand fabrics using Unified Fabric Manager (UFM).
  • Configure and operate NVIDIA BlueField DPUs for networking, security, and storage offload.
  • Optimize data pipelines between storage and GPU infrastructure.

Security & Compliance

  • Apply CKS-level security practices to Kubernetes clusters and AI workloads.
  • Implement service mesh, microsegmentation, and secure workload architectures.
  • Conduct vulnerability assessments, audits, and policy enforcement.

Automation, Observability & MLOps

  • Automate infrastructure provisioning using Terraform and Ansible.
  • Develop automation using Python, Bash, and YAML.
  • Monitor AI infrastructure using NVIDIA DCGM, Prometheus, and Grafana.
  • Build MLOps workflows using Kubeflow Pipelines and NVIDIA Triton Inference Server.
  • Troubleshoot complex issues across hardware, networking, Kubernetes, and AI software layers.

Required Skills & Experience

  • Hands-on experience running AI/ML workloads on NVIDIA DGX systems.
  • Strong expertise in Kubernetes administration, architecture, and security.
  • Deep experience with InfiniBand, UFM, and BlueField DPU administration.
  • Strong scripting and automation skills with Python, Bash, and YAML.
  • Experience designing scalable, secure, production-grade AI infrastructure.
  • CKA, CKAD, and CKS certifications.

Preferred Qualifications

  • Experience with NVIDIA Base Command Manager.
  • Experience with NVIDIA GPU Operator.
  • Experience with Kubeflow/Kubeflow Pipelines.
  • Knowledge of MIG GPU partitioning.
  • Experience with Terraform and Ansible.
  • Experience with NVIDIA Triton Inference Server.
  • Strong collaboration skills with ML researchers, DevOps engineers, and infrastructure teams.

Technical Environment

GPU Infrastructure: NVIDIA DGX, BasePOD, SuperPOD, NVIDIA AI Enterprise
Kubernetes: Kubernetes, GPU Operator, Helm, Custom Controllers, MIG
Networking: InfiniBand, UFM, BlueField DPU
MLOps: Kubeflow Pipelines, NVIDIA Triton Inference Server
Automation: Terraform, Ansible, Python, Bash, YAML
Monitoring: NVIDIA DCGM, Prometheus, Grafana

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10371609
  • Position Id: 9064158
  • Posted 5 hours ago

Company Info

About Prabhav Services Inc

High ROI

Many companies find that constant maintenance eats into their budget for new technology. By outsourcing your IT management to us, you can focus on what you do best--running your business.

Satisfaction Guaranteed

That's why our goal is to provide an experience that is tailored to your company's needs. No matter the budget, we pride ourselves on providing professional customer service.

Technical Experience

We are well-versed in a variety of operating systems, networks, and databases. We use this expertise to help our customers with a variety of small to mid-sized projects.

Contact the job poster
AS

Aniket Sharma

Recruiter @ Prabhav Services Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Easy Apply

Contract

Depends on Experience

Remote

Today

Easy Apply

Contract

Depends on Experience

Remote

Today

Easy Apply

Contract

Depends on Experience

Search all similar jobs