Infrastructure Engineer (Kubernetes, GPU/NVIDIA, Terraform)

St. Louis, MO, US • Posted 5 days ago • Updated 5 days ago
Contract W2
1 Month
No Travel Required
On-site
$70 - $80/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Kubernetes
  • NVIDIA
  • GPU
  • Terraform
  • Infrastructure

Summary

Infrastructure Engineer (Kubernetes, GPU/NVIDIA, Terraform)

Introduction

We are seeking a highly skilled AI Infrastructure Engineer to design, build, and support scalable GPU-enabled Kubernetes platforms for AI/ML workloads. The ideal candidate will have strong expertise in Kubernetes, Linux, Terraform, and NVIDIA GPU technologies, with experience building enterprise-grade AI infrastructure.

Responsibilities

  • Design, deploy, and manage production-grade Kubernetes clusters
  • Build and maintain GPU-enabled infrastructure for AI/ML workloads
  • Deploy and manage Longhorn storage solutions
  • Automate infrastructure provisioning using Terraform
  • Monitor and optimize platform performance using Prometheus and Grafana
  • Troubleshoot Kubernetes, GPU, storage, and infrastructure issues
  • Collaborate with ML engineers, data scientists, and DevOps teams
  • Create technical documentation, runbooks, and operational guides
  • Conduct knowledge transfer (KT) sessions and hands-on workshops for client teams
  • Provide guidance on Kubernetes operations, NVIDIA GPU management, Longhorn storage, and Canonical ecosystem tools
  • Advise on cloud-native infrastructure, AI/ML platform best practices, scalability, reliability, and cost optimization
  • Ensure successful production handoff with full operational readiness

Requirements

Must-Have Skills

  • Strong hands-on experience with Kubernetes
  • Expertise in Linux (Ubuntu)
  • Experience with Terraform (Infrastructure as Code)
  • Experience with GPUs and NVIDIA Stack (CUDA, NVIDIA GPU Operator, device plugins, scheduling)
  • Strong troubleshooting and automation skills

Preferred Skills

  • Experience with Longhorn or Ceph storage
  • Knowledge of Canonical ecosystem (MAAS, Juju, Charmed Kubernetes)
  • Experience with Kubeflow or other AI/ML platform tools
  • Certifications such as CKA or NVIDIA Certified

Required Skills

  • Kubernetes
  • Linux (Ubuntu)
  • Terraform
  • NVIDIA GPU Stack
  • GPU Infrastructure
  • Prometheus
  • Grafana
  • AI Infrastructure
  • Infrastructure Automation

Nice to Have

  • Longhorn
  • Ceph
  • Kubeflow
  • MAAS
  • Juju
  • Charmed Kubernetes
  • CKA Certification
  • NVIDIA Certification
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10336460
  • Position Id: 9046438
  • Posted 5 days ago
Contact the job poster
DR

Devendranath Reddipalli

Recruiter @ ITCAPS LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

St. Louis, Missouri

Today

Easy Apply

Contract, Third Party

$DOE

Remote

Today

Easy Apply

Contract

$55 - $65 per hour

Remote

Today

Full-time

No location provided

Today

Full-time

USD 111,000.00 - 231,250.00 per year

Search all similar jobs