Principal Architect GPU Platform & Orchestration

Hybrid in Plano, TX, US • Posted 1 day ago • Updated 1 day ago
Contract Independent
12 Months
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • GPUaaS
  • GPU Platform
  • Kubernetes
  • OpenShift
  • NVIDIA GPU Operator
  • NVIDIA Network Operator
  • GPU Device Plugin
  • MIG
  • MIG Manager
  • Node Feature Discovery
  • Multi-Tenancy
  • RBAC
  • Network Policies
  • GPU Scheduling
  • Kueue
  • Volcano
  • Slurm
  • HPC
  • Terraform
  • Helm
  • GitOps
  • Argo CD
  • Flux
  • Ansible
  • DCGM
  • GPU Metering
  • Platform Architecture
  • Offshore Team Leadership
  • Customer-Facing Architecture

Summary

Job Title: Principal Architect GPU Platform & Orchestration
Client: LTTS
Location: Plano, TX Hybrid Employment Type: FTE
Rate: $/hr on 1099 Experience: 7+ Years Platform Engineering; 4+ Years Kubernetes in Production Interview Mode: Virtual Practice: AI Infrastructure / GPU-as-a-Service

About the Role
We are building a GPU-as-a-Service and AI factory practice from the ground up, delivering multi
tenant GPU platforms for enterprise and industrial customers. This is the senior technical seat on
that platform. You will own the orchestration and multi-tenancy architecture that turns a GPU
cluster into a consumable service, set the standards our global delivery team builds against, and
serve as deputy to the practice lead in customer architecture engagements.
This is an architecture role. You will design, review, and defend and lead an offshore
engineering pod that executes.
What you ll do

Architect the GPUaaS control plane on Kubernetes and OpenShift: NVIDIA GPU Operator,
Network Operator, device plugin, MIG manager, node feature discovery.
Design multi-tenancy end to end MIG partitioning strategy, time-slicing tiers, namespace
and RBAC model, network policy, quotas, priority classes, and tenant onboarding.
Own GPU scheduling and allocation policy: gang scheduling (Kueue, Volcano), fair-share
and preemption, topology-aware placement, and Slurm integration where customers run
genuine batch HPC.
Define the service catalog instance shapes, self-service request flow, and GPU metering
for chargeback or showback from DCGM telemetry.
Build and own the reusable platform blueprint: reference architecture, Terraform and
Helm modules, GitOps patterns, and runbooks that every engagement starts from.
Technically lead an offshore delivery pod set standards, run design reviews, gate
deliverables before they reach a customer.
Partner with the practice lead on customer discovery, solution design, and technical
escalation; lead design sessions independently as the practice scales.
What you need
7+ years platform engineering, with 4+ on Kubernetes in production; OpenShift experience
valued.
Demonstrated GPU workload orchestration GPU Operator, MIG, device plugin, GPU
scheduling policy on real multi-node clusters.
Real multi-tenancy design experience: isolation, quota, RBAC, network segmentation, and
the failure modes each produces.
Batch or HPC scheduling background (Slurm, LSF, PBS) or gang scheduling on Kubernetes.
Strong IaC and GitOps: Terraform, Helm, Argo CD or Flux, Ansible.
Experience leading distributed or offshore engineering teams through written standards
rather than direct supervision.
Customer-facing credibility you can whiteboard a design for a CTO and defend it under
challenge.
Nice to have
NVIDIA AI Enterprise; Run:ai or equivalent GPU orchestration; consulting or professional-services
background; internal developer platform / service catalog experience; CKA or CKS.
First 90 days
v1 GPUaaS platform blueprint published and deployed in our reference environment; tenancy and
metering model validated; leading a customer design session unaccompanied; offshore pod
onboarded to your standards.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10202400
  • Position Id: 9088286
  • Posted 1 day ago
Contact the job poster
DG

Dolly Gupta

Recruiter @ VST Consulting, Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Dallas, Texas

9d ago

Easy Apply

Full-time

Depends on Experience

Hybrid in Irving, Texas

Yesterday

Easy Apply

Contract, Third Party

$66.5

Dallas, Texas

Today

Full-time

USD 97,725.00 per year

Irving, Texas

Today

Full-time

USD 60.00 - 71.00 per hour

Search all similar jobs