π¨ π₯ NEW PRIORITY HIRING | 10 HPC & Kubernetes Positions | Amazon/AWS
We are actively hiring for 10 high-priority HPC & Kubernetes roles supporting Amazon/AWS!
πΌ Employment: FTE ONLY β NO CONTRACTORS
π Location: 100% Remote β US
π° Total Cash Compensation: $115K
π The Opportunity
Our AWS customer develops and manages large-scale HPC clusters across AWS, CoreWeave, Google Cloud Platform, and other cloud providers.
The environment currently supports several thousand GPUs, with plans to scale to 10X in 2026 and beyond.
This is a Kubernetes-heavy infrastructure role where reliability and scale are critical. Misconfigurations or failed upgrades can directly impact thousands of GPU-hours, and the environment regularly presents complex, novel infrastructure challenges.
πΉ Key Responsibilities
-
Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across multiple cloud providers.
-
Own cluster lifecycle, node pool management, networking policies, and platform stability.
-
Provision HPC infrastructure through CI/CD across AWS, CoreWeave, Google Cloud Platform, OCI, and additional providers.
-
Manage job scheduling and optimize GPU compute allocation for training and inference workloads.
-
Define and maintain SLIs/SLOs and build monitoring and alerting solutions.
-
Participate in severity escalation and incident response.
-
Author detailed post-incident reviews and drive continuous improvement.
-
Partner daily with Networking, Storage, Security, and AI/ML Platform teams.
β
Required Qualifications
-
4+ years of experience in infrastructure engineering, cloud platforms, or HPC.
-
Strong hands-on Kubernetes experience operating clusters at meaningful scale.
-
Experience with:
-
Terraform proficiency β daily infrastructure-as-code development and review.
-
Working knowledge of AWS, including:
-
EC2
-
S3
-
EFS
-
FSx for Lustre
-
Strong Python skills for tooling and automation.
-
HPC and/or GPU infrastructure experience strongly preferred.
β οΈ Candidates with Kubernetes experience limited to small, local, or lab environments are unlikely to be a fit.
π REQUIRED SELF-RATING
Please include a 1β5 self-rating for each skill when submitting your profile:
| Skill |
Self-Rating (1β5) |
| Kubernetes (EKS, CKS, GKE) |
β |
| HPC Infrastructure |
β |
| GPU Infrastructure |
β |
| Terraform |
β |
| AWS (EC2, S3, EFS, FSx for Lustre) |
β |
| Python |