PERM Role-Direct Client-HPC Systems Engineer (High Performance Computing)

Dallas, TX, US • Posted 19 hours ago • Updated 19 hours ago
Full Time
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • HPC
  • high performance computing
  • GPU
  • NVIDIA ecosystem
  • Kubernetes
  • plugins
  • python
  • observability
  • prometheus
  • grafana
  • RBAC
  • CI/CD

Summary

NOTE-THIS IS A PERM ROLE WITH MY DIRECT CLIENT IN DALLAS AND ONSITE FROM DAY 1. NO 3RD PARTIES, NO VISA CANDIDATES.

Role: Sr. Kubernetes Engineer (HPC, GPU, NVIDIA) Job Location: Dallas, TX Duration: Permanent Hire/Direct Hire Work Model: Onsite Interview: MS Teams Video Education: Bachelor's Degree

Requirements:

  • Strong experience using Kubernetes in production environments.
  • Experience working with NVIDIA GPUs and Kubernetes.
  • Experience with NVIDIA GPU Operator, device plugins, NVML, MIG, and DCGM.
  • Proficiency in Go or Python for developing Kubernetes operators and controllers.
  • Strong understanding of Kubernetes internals, including CRDs, RBAC, custom controllers, and scheduler extensions.
  • Experience supporting GPU-intensive workloads such as large language models (LLMs), training pipelines, and scientific computing.
  • Hands-on experience with Helm, Kustomize, and GitOps workflows.
  • Familiarity with CNI plugins, especially NVIDIA CNI and Multus.
  • Experience monitoring GPU metrics and cluster health using Prometheus and DCGM Exporter.

Responsibilities:

  • Architecting and operating Kubernetes clusters optimized for GPU workloads, leveraging NVIDIA GPU Operator, Network Operator and DCGM.
  • Developing, deploying and maintaining custom Kubernetes operators and controllers to automate infrastructure services.
  • Integrating NVIDIA device plugins, Multi-Instance GPU (MIG) and GPU sharing features into the scheduling layer.
  • Collaborating with HPC, ML and DevOps teams to ensure multi-tenant, high-throughput cluster performance.
  • Driving observability and telemetry integrations using Prometheus, Grafana, DCGM Exporter and OpenTelemetry.
  • Implementing secure multi-user and multi-namespace GPU isolation, with RBAC and policy enforcement, such as OPA or Gatekeeper.
  • Maintaining CI/CD pipelines for Kubernetes infrastructure using GitOps, ArgoCD and FluxCD.
  • Contributing to infrastructure-as-code, using Terraform, Helm, and Kustomize.

Share resume at

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10212185
  • Position Id: 9078990
  • Posted 19 hours ago

Company Info

About The Brixton Group

We are a values-based professional technology services firm with over 150 years of combined experience in the technology industry. We re deeply engaged in the successful pairing of the right people to the right projects, and we attribute our national presence to the referrals we ve received through current and past clients and candidates.



Our History

-Established in 1998

-Woman-Owned Business with national footprint



AWARDS

-Four straight years listed on Inc. Magazine s Fastest-Growing Private Companies in America.

-Three straight years listed on Charlotte s Fast 50.



We believe that when people find the ideal setting to express their talents, the possibilities are infinite. This is why we exist. We are inspired and guided by a greater purpose than profit.


Career Opportunities
Contact the job poster
Nagesh Rao

Nagesh Rao

Sr. Executive Resourcing @ The Brixton Group
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs