
Remote
•
Today
The Work AWS customer develops and manages several HPC clusters across AWS, CoreWeave, Google Cloud Platform, and other providers. Several thousand GPUs today scaling to 10X in 2026 and beyond. This role is Kubernetes-heavy. You'll operate multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours. The clusters are large enough that novel failure modes are routine. Responsibilities Operate Kubernetes platforms (EK
Easy Apply
Full-time, Third Party
Depends on Experience
















