Lead HPC Kubernetes Engineer

Remote • Posted 12 hours ago • Updated 12 hours ago
Full Time
Remote
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • ProVision
  • Continuous Integration
  • Continuous Delivery
  • OCI
  • Management
  • Job Scheduling
  • GPU
  • Training
  • Computer Networking
  • Storage
  • Artificial Intelligence
  • Machine Learning (ML)
  • Cloud Computing
  • HPC
  • Debugging
  • Terraform
  • Writing
  • Amazon EC2
  • Amazon S3
  • Amazon EFS
  • Python
  • English
  • Amazon Web Services
  • Kubernetes
  • Google Cloud Platform
  • Google Cloud

Summary

We are seeking a Lead HPC Kubernetes Engineer to help our customer develop and manage several HPC clusters across AWS, CoreWeave, Google Cloud Platform, and other providers, spanning several thousand GPUs today and scaling to 10x in 2026 and beyond. This role is Kubernetes-heavy, operating multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours, at a scale where novel failure modes are routine. Responsibilities Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers Take ownership of cluster lifecycle, node pool management, networking policy, and stability maintenance during rapid growth Provision HPC infrastructure through CI/CD systems across AWS, CoreWeave, Google Cloud Platform, and OCI, with additional providers to be added in the near future Manage job scheduling to allocate GPU compute across training and inference workloads Define and maintain SLIs/SLOs Build monitoring and alerting systems Participate in severity escalation response and author post-incident reviews Coordinate daily with Networking, Storage, Security, and AI/ML platform teams Requirements 5+ years of experience in infrastructure engineering, cloud platforms, or HPC Expertise in Kubernetes, with hands-on experience operating clusters at meaningful scale, including node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets Proficiency in Terraform for writing and reviewing infrastructure-as-code daily Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre) Skills in Python for tooling and automation English proficiency at B2 level or higher Nice to have Familiarity with Google Kubernetes Engine Familiarity with Amazon Elastic Kubernetes Service Knowledge of Google Cloud Platform
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10330481
  • Position Id: 5a52f8a1a486a1288b6d024ea5b65c18
  • Posted 12 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

•

2d ago

Easy Apply

Full-time

$100,000 - $120,000

Remote

•

2d ago

Easy Apply

Full-time

130,000 - 140,000

Remote

•

2d ago

Easy Apply

Full-time, Third Party

$120,000 - $140,000

Remote or Texas City, Texas

•

Today

Full-time

Search all similar jobs