Senior AI Infrastructure Engineer – NVIDIA GPU / GB300

Remote • Posted 1 hour ago • Updated 1 hour ago
Contract W2
6 Months
Occasional Travel Required
Remote
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Artificial Intelligence
  • Ansible
  • Authorization
  • BMC
  • BIOS
  • Cloud Computing
  • CUDA
  • Computer Cluster Management
  • Bash
  • Computer Hardware
  • Computer Networking
  • Employment Authorization
  • Enterprise Networks
  • Ethernet
  • File Systems
  • Firmware
  • GC
  • GPU
  • Incident Management
  • Ganglia
  • Linux Administration
  • Kubernetes
  • Integrated Circuit
  • Infrastructure Lifecycle Management
  • HPC
  • Lifecycle Management
  • Grafana
  • Management
  • Network Engineering
  • IPMI
  • InfiniBand
  • Machine Learning (ML)
  • Optimization
  • Orchestration
  • Workflow
  • Storage
  • Spectrum
  • Remote Direct Memory Access
  • Resource Management
  • Production Support
  • Provisioning
  • Python
  • Root Cause Analysis
  • Scheduling
  • Scripting

Summary

JOB TITLE: Senior AI Infrastructure Engineer – NVIDIA GPU / GB300

JOB TYPE: Contract

LOCATION: Remote – United States

TRAVEL: Must be willing to travel to the Bay Area / San Jose, CA as required

JOB DESCRIPTION

We are seeking a Senior AI Infrastructure Engineer with strong hands-on experience deploying, configuring, and operating large-scale NVIDIA GPU and HPC infrastructure.

The ideal candidate will have direct experience with NVIDIA GB200, GB300, B200, B300, Blackwell, HGX, NVL72, or comparable GPU platforms. This role requires hands-on experience with rack-scale GPU deployments, cluster provisioning, AI/HPC workload orchestration, high-performance networking, firmware, monitoring, and infrastructure automation.

The successful candidate should be comfortable working across compute, networking, storage, cooling, cluster management, and data center infrastructure.

KEY RESPONSIBILITIES

• Deploy and bring up large-scale NVIDIA GPU clusters and AI infrastructure.

• Perform rack-scale GPU infrastructure integration, provisioning, configuration, validation, and production readiness activities.

• Configure and manage GPU clusters using NVIDIA Base Command Manager or similar cluster management platforms.

• Perform BIOS, BMC, IPMI, Redfish, firmware, and hardware validation activities.

• Support GPU cluster lifecycle management, node provisioning, health monitoring, and troubleshooting.

• Configure and troubleshoot high-performance GPU networking environments including InfiniBand, NVLink, NVSwitch, RoCE, and high-speed Ethernet.

• Support NVIDIA networking platforms, GPU fabrics, and large-scale AI/HPC cluster connectivity.

• Deploy and manage Kubernetes-based GPU environments for AI/ML workloads.

• Support Slurm-based HPC workload scheduling, resource management, queue configuration, and workload optimization.

• Automate infrastructure provisioning, validation, monitoring, and operational workflows using Python, Ansible, Bash, or similar tools.

• Monitor and troubleshoot GPU cluster performance, networking, compute, storage, and infrastructure issues.

• Work with engineering, data center, networking, facilities, and hardware teams to resolve deployment and operational issues.

• Support telemetry, observability, incident response, and infrastructure reliability initiatives.

• Participate in hardware/firmware upgrades, system validation, root-cause analysis, and production support.

REQUIRED SKILLS

• Strong hands-on experience with NVIDIA GPU infrastructure.

• Experience with NVIDIA GB200 / GB300, B200 / B300, Blackwell, HGX, NVL72, or similar GPU platforms.

• Experience with large-scale GPU cluster or AI Factory deployments.

• Experience with GPU provisioning, cluster bring-up, and infrastructure lifecycle management.

• Experience with NVIDIA Base Command Manager or similar HPC cluster management tools.

• Strong Linux administration and troubleshooting skills.

• Experience with Kubernetes and/or Slurm.

• Experience with InfiniBand and/or high-performance Ethernet/RoCE.

• Experience with NVLink / NVSwitch or GPU fabric technologies.

• Experience with firmware, BIOS, BMC, IPMI, Redfish, and hardware validation.

• Strong scripting/automation experience using Python, Bash, Ansible, or similar technologies.

• Experience supporting HPC, AI/ML, or large-scale GPU workloads.

PREFERRED EXPERIENCE

• NVIDIA GB200 / GB300 NVL72

• NVIDIA Blackwell architecture

• DGX SuperPOD

• NVIDIA Spectrum-X

• NVIDIA BlueField DPU

• GPUDirect / RDMA

• CUDA

• Run:ai

• DDN / Lustre / parallel file systems

• Prometheus / Grafana / Ganglia

• Direct-to-chip liquid cooling and GPU data center infrastructure

• Large-scale AI Factory deployments

CANDIDATE PROFILE

We are looking for candidates with demonstrated hands-on experience rather than candidates who have only worked with these technologies at a high level.

Candidates should be able to explain:

• GPU platforms and clusters they have personally deployed or supported

• Number of GPUs, racks, or clusters involved

• Their specific responsibilities during deployment and operations

• Networking technologies used

• Kubernetes / Slurm implementation

• Firmware, provisioning, monitoring, and troubleshooting responsibilities

• AI/HPC workloads supported

LOCATION & TRAVEL

This is a remote position within the United States.

Candidates must be willing and able to travel to the Bay Area / San Jose, California, as required.

Please include the candidate's current location and willingness to travel with each submission.

WORK AUTHORIZATION

SUBMISSION REQUIREMENTS

Please provide the following with each candidate submission:

• Updated Resume

• Current Location

• Work Authorization

• Earliest Availability

• Willingness to Travel to Bay Area / San Jose, CA

• Brief summary of relevant NVIDIA GPU / AI infrastructure experience

IMPORTANT

Please do not submit candidates whose experience is primarily limited to:

• Generic enterprise networking

• Traditional network engineering

• Generic cloud engineering

• Basic Kubernetes administration

• Generic data center operations

• General Linux administration without GPU/HPC experience

• General project management

Candidates with direct NVIDIA GPU, AI Factory, HPC, GB200/GB300, Blackwell, InfiniBand, NVLink/NVSwitch, Kubernetes, Slurm, and large-scale cluster deployment experience will be strongly preferred.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91172187
  • Position Id: 9108775
  • Posted 1 hour ago

Company Info

About Texnere Americas Inc

At Texnere, we are committed to helping businesses find the right talent, at the right time, through the right approach. Our 360° talent solutions seamlessly integrate flexible hiring models—Flexible Staffing, Team Leasing, Synergizing Agenting AI, and Managed Services—with deep industry expertise across IT, BPM, Sales & Marketing, and more.

We understand that every organization is unique, operating within specific sectors and organizational ‘vectors’ that present distinct challenges and opportunities. Whether you're a healthcare startup looking to innovate or a global captive center (GCC) in the tech space, Texnere provides specialized talent solutions tailored to your precise needs.

About_Company_OneAbout_Company_Two
Contact the job poster
PV

Peddinti Venkat Moudgalya

Recruiter @ Texnere Americas Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs