HPC & Kubernetes Positions

Remote β€’ Posted 15 hours ago β€’ Updated 15 hours ago
Full Time
No Travel Required
Remote
$150 - $155/hr
Fitment

Dice Job Match Scoreβ„’

πŸ”’ Crunching numbers...

Job Details

Skills

  • Operate Kubernetes platforms (EKS
  • CKS
  • GKE) at significant scale across multiple cloud providers. Own cluster lifecycle
  • node pool management
  • networking policies
  • and platform stability. Provision HPC infrastructure through CI/CD across AWS
  • CoreWeave
  • GCP
  • OCI
  • and additional providers. Manage job scheduling and optimize GPU compute allocation for training and inference workloads. Define and maintain SLIs/SLOs and build monitoring and alerting solutions. Participate in severity escalation and incident response. Author detailed post-incident reviews and drive continuous improvement. Partner daily with Networking
  • Storage
  • Security
  • and AI/ML Platform teams.

Summary

🚨 πŸ”₯ NEW PRIORITY HIRING | 10 HPC & Kubernetes Positions | Amazon/AWS

We are actively hiring for 10 high-priority HPC & Kubernetes roles supporting Amazon/AWS!


πŸ’Ό Employment: FTE ONLY β€” NO CONTRACTORS
🌎 Location: 100% Remote β€” US
πŸ’° Total Cash Compensation: $115K

πŸš€ The Opportunity

Our AWS customer develops and manages large-scale HPC clusters across AWS, CoreWeave, Google Cloud Platform, and other cloud providers.

The environment currently supports several thousand GPUs, with plans to scale to 10X in 2026 and beyond.

This is a Kubernetes-heavy infrastructure role where reliability and scale are critical. Misconfigurations or failed upgrades can directly impact thousands of GPU-hours, and the environment regularly presents complex, novel infrastructure challenges.

πŸ”Ή Key Responsibilities

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across multiple cloud providers.

  • Own cluster lifecycle, node pool management, networking policies, and platform stability.

  • Provision HPC infrastructure through CI/CD across AWS, CoreWeave, Google Cloud Platform, OCI, and additional providers.

  • Manage job scheduling and optimize GPU compute allocation for training and inference workloads.

  • Define and maintain SLIs/SLOs and build monitoring and alerting solutions.

  • Participate in severity escalation and incident response.

  • Author detailed post-incident reviews and drive continuous improvement.

  • Partner daily with Networking, Storage, Security, and AI/ML Platform teams.

βœ… Required Qualifications

  • 4+ years of experience in infrastructure engineering, cloud platforms, or HPC.

  • Strong hands-on Kubernetes experience operating clusters at meaningful scale.

  • Experience with:

    • Node pool sizing

    • Scheduler troubleshooting

    • CNI troubleshooting

    • Large-scale rolling upgrades

  • Terraform proficiency β€” daily infrastructure-as-code development and review.

  • Working knowledge of AWS, including:

    • EC2

    • S3

    • EFS

    • FSx for Lustre

  • Strong Python skills for tooling and automation.

  • HPC and/or GPU infrastructure experience strongly preferred.

⚠️ Candidates with Kubernetes experience limited to small, local, or lab environments are unlikely to be a fit.

πŸ“Š REQUIRED SELF-RATING

Please include a 1–5 self-rating for each skill when submitting your profile:

Skill Self-Rating (1–5)
Kubernetes (EKS, CKS, GKE) ⭐
HPC Infrastructure ⭐
GPU Infrastructure ⭐
Terraform ⭐
AWS (EC2, S3, EFS, FSx for Lustre) ⭐
Python
Employers have access to artificial intelligence language tools (β€œAI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: RTX17ee7d
  • Position Id: 1671-33630-1791404932
  • Posted 15 hours ago
Contact the job poster
BU

Batch User

Recruiter @ SoftPath Technologies LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

β€’

Today

Easy Apply

Full-time, Third Party

Depends on Experience

Remote

β€’

Today

Easy Apply

Full-time, Third Party

Depends on Experience

Remote

β€’

2d ago

Easy Apply

Full-time

$180,000 - $200,000

Remote

β€’

Today

Easy Apply

Full-time

Depends on Experience

Search all similar jobs