HPC-High Performance Computing Consultant

Remote • Posted 18 hours ago • Updated 18 hours ago
Full Time
Occasional Travel Required
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • HPC
  • High-Performance Computing
  • Slurm
  • python
  • Kubernetes

Summary

Role: HPC (High-Performance Computing) Consultant (5+ Openings)

Location: Remote

Duration: Fulltime 

Must Have: Kubernetes + Slurm + NVIDIA GPU + AI/ML Infrastructure + Terraform + Python + AWS (EKS/FSx Lustre) + Monitoring/SRE

 

We are looking for engineers with expertise across Kubernetes, cloud infrastructure, HPC platforms, GPU computing, Terraform, and automation. Depending on experience, candidates may be considered for Platform Engineer, Kubernetes Engineer, HPC Engineer, Cloud Infrastructure Engineer, DevOps Engineer, AI Infrastructure Engineer, or Site Reliability Engineer (SRE) roles.

 

Key Responsibilities

  • Design, deploy, operate, and support large-scale Kubernetes platforms across AWS, Google Cloud Platform, CoreWeave, OCI, and other cloud environments.
  • Manage Kubernetes cluster lifecycle activities including provisioning, scaling, node pool management, upgrades, troubleshooting, and performance optimization.
  • Support AI/ML and HPC workloads, including GPU-enabled compute infrastructure for training and inference environments.
  • Provision and automate cloud and infrastructure resources using Terraform and CI/CD pipelines.
  • Troubleshoot Kubernetes scheduling, networking, storage, and platform reliability issues.
  • Implement monitoring, observability, alerting, SLIs/SLOs, and incident response processes.
  • Collaborate with Networking, Security, Storage, AI/ML, Data Engineering, and Application teams.
  • Develop automation and operational tooling using Python and cloud-native technologies.
  • Participate in production support, root cause analysis, capacity planning, and platform optimization initiatives.

 

Required Skills

Kubernetes & Container Platforms

  • Kubernetes (EKS, GKE, AKS, OpenShift, CoreWeave CKS)
  • Cluster Lifecycle Management
  • Node Pool Management
  • Scheduler Troubleshooting
  • CNI Troubleshooting
  • Networking Policies
  • RBAC
  • Helm
  • Autoscaling
  • Rolling Upgrades

 

Cloud Infrastructure

  • AWS (EC2, S3, IAM, VPC, EKS, EFS, FSx for Lustre)
  • Google Cloud Platform (Google Cloud Platform)
  • OCI (Oracle Cloud Infrastructure)
  • Azure (Preferred)
  • Multi-Cloud Infrastructure

 

Infrastructure as Code & Automation

  • Terraform
  • Infrastructure as Code (IaC)
  • CI/CD Pipelines
  • GitHub Actions / Jenkins / GitLab CI
  • Ansible (Preferred)

 

Programming & Scripting

  • Python
  • Bash/Shell Scripting
  • Automation Development
  • REST API Integration

 

HPC & GPU Infrastructure (Preferred)

  • High-Performance Computing (HPC)
  • Slurm
  • NVIDIA GPU Platforms
  • GPU Scheduling
  • CUDA
  • Distributed Computing
  • AI/ML Infrastructure

 

Monitoring & Reliability

  • Prometheus
  • Grafana
  • Datadog
  • Splunk
  • Cloud Monitoring
  • SLI/SLO Management
  • Incident Response
  • Root Cause Analysis (RCA)

 

Preferred Experience

Experience with any of the following is highly desirable:

  • AI/ML Infrastructure Platforms
  • Kubeflow
  • KServe
  • Ray
  • MLflow
  • vLLM
  • Vector Databases
  • Distributed Training Platforms
  • CoreWeave
  • AWS ParallelCluster
  • FSx for Lustre
  • Lustre
  • WekaFS
  • InfiniBand
  • RDMA
  • AI Training & Inference Workloads
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10111282
  • Position Id: 9109171
  • Posted 18 hours ago
Create job alert
Never miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

•

Yesterday

Easy Apply

Full-time, Third Party

$150 - $155

Remote

•

3d ago

Easy Apply

Full-time

$180,000 - $200,000

Remote

•

Yesterday

Easy Apply

Full-time, Third Party

Depends on Experience

Remote or Armonk, New York

•

Today

Contract

$90 - $100 hourly

Search all similar jobs