NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems)

Remote • Posted 1 hour ago • Updated 1 hour ago
Contract W2
Contract Independent
Remote
$90 - $100/hr
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Nvidia DGX
  • GPU Cluster
  • Kubernetes

Summary

NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems) Department: Infrastructure Engineering

Location / Remote Policy: Remote Role Type: Contract 6-month initial engagement

About Our Client
Our client is a technology and professional services firm founded in 2015 on the strength of its founders' 30 years of industry experience. They set out to bridge a gap in professional services to be a true partner rather than just a vendor delivering expert guidance, innovative solutions, and personalized service at a cost-effective rate. Their mission is to empower businesses to succeed in the digital era, harnessing technology to drive transformation, innovation, and growth. Guided by a "make a customer, not a sale" philosophy, they lead with a customer-first approach and a team of senior-level engineers sourced from the world's leading OEMs, including AWS, Palo Alto Networks, Cisco, and Microsoft.

Job Description
Our client is seeking a highly skilled AI Infrastructure & Kubernetes Platform Engineer with a proven track record deploying and managing NVIDIA DGX-based AI clusters, orchestrating containerized AI workloads on Kubernetes, and ensuring secure, high-throughput operations across InfiniBand-powered networks. You'll bring a strong certification foundation across both Kubernetes (CKA, CKAD, CKS) and NVIDIA's AI infrastructure stack, paired with hands-on experience across DGX, BlueField, and high-speed networking.
This role is central to supporting AI/ML infrastructure at scale enabling efficient training and inference for complex models and integrating NVIDIA's compute, storage, and fabric solutions with modern DevOps practices. Day to day, you'll own DGX cluster operations, architect GPU-accelerated Kubernetes platforms, tune InfiniBand fabric for throughput, and harden the environment through DPU-enhanced security.
You'll work at the intersection of infrastructure, DevOps, and AI/ML, keeping the platform reliable and cost-efficient for the teams that depend on it. The ideal candidate is deeply hands-on, obsessed with performance and security, and energized by operating some of the most advanced AI compute available.

Duties and Responsibilities
AI Infrastructure Operations

  • Deploy and manage NVIDIA DGX BasePODs and SuperPODs for high-performance AI workloads.
  • Oversee DGX system lifecycle operations, including provisioning, monitoring, firmware upgrades, and capacity planning.
  • Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools.
  • Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification after deployments or hardware changes.

Kubernetes Platform Engineering

  • Architect secure, scalable Kubernetes clusters optimized for GPU-accelerated workloads using the NVIDIA GPU Operator.
  • Apply CKA/CKAD/CKS expertise to develop, deploy, and secure AI applications on Kubernetes.
  • Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows.

High-Performance Networking & DPUs

  • Administer InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM).
  • Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput.
  • Use BlueField to offload storage, firewalling, and telemetry, strengthening AI workload security and performance.

Security & Compliance

  • Apply CKS best practices to secure containerized AI environments.
  • Configure runtime security, secrets management, network segmentation, and auditing across DPU-enhanced Kubernetes deployments.
  • Support zero-trust initiatives by enforcing workload identity, RBAC policies, and supply-chain integrity across AI container images and model artifacts.

Monitoring, Telemetry & Optimization

  • Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs.
  • Tune system performance and model-training pipelines for cost-efficiency and throughput.
  • Build and maintain operational runbooks, incident-response playbooks, and SLA dashboards covering GPU utilization, thermal thresholds, and fabric health.

Required Experience/Skills
Certifications

  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD)
  • Certified Kubernetes Security Specialist (CKS)
  • NVIDIA Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
  • NVIDIA Certified Professional: AI Infrastructure (NCP-AII)
  • NVIDIA Certified Professional: AI Operations (NCP-AIO)
  • NVIDIA Certified Professional: AI Networking (NCP-AIN)

Hands-On Expertise

  • DGX System, BasePOD, and SuperPOD administration
  • BlueField DPU configuration and operations
  • InfiniBand fabric and UFM management
  • Base Command Manager for workload orchestration

Technical Skills

  • Kubernetes, Helm, and the NVIDIA GPU Operator
  • DevOps tooling: Ansible, Terraform, GitOps, CI/CD pipelines
  • Programming/scripting: Python, YAML, Bash

Nice-to-Haves

  • Kubeflow and broader MLOps pipeline experience.
  • Parallel/HPC storage: NFS, BeeGFS, Lustre.
  • Advanced networking: RoCE, RDMA, gRPC, and DPU offload tuning.

Education
Bachelor's degree in Computer Science, Engineering, or a related field or equivalent hands-on experience.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91138303
  • Position Id: 9060740
  • Posted 1 hour ago

Company Info

About One IT Corp

One IT Corp was founded with the vision of providing advanced IT solutions to our growing customer base. The company adopts high commercial values of transparency and integrity in the exercise of its activities. One IT Corp provides exemplary services through innovation, technical expertise, and fair business practices.

Over the years, One IT Corp has defined, designed and developed business solutions based on technology and processes that help its customers differentiate themselves from others. Focusing on one of the objectives and encouragement for Business Intelligence (BI) tools, application development, systems integration, software development, testing, recruitment, and training in the company, who have created milestones throughout the process.

One IT Corp is proud to build long-term relationships with its customers. We proudly emphasize that our only motto is the delight of the customer, achieved through exemplary service and respect for values and standards.

About_Company_OneAbout_Company_Two
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs