AI GPU Cluster Operations Engineer

Barker, NY, US • Posted 5 days ago • Updated 5 days ago
Full Time
No Travel Required
Able to Sponsor
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Server Hardware
  • Computer Hardware
  • Linux
  • GPU
  • InfiniBand
  • RMA
  • Root Cause Analysis
  • Technical Support
  • Mandarin

Summary

Position Overview

Responsible for the operations and maintenance of large-scale AI GPU computing clusters, including NVIDIA, AMD and/or other GPU server platforms, as well as associated Data Center Network (DCN), Out-of-Band Management (OOB) network, and high-performance GPU fabric, including InfiniBand and RDMA Ethernet.

The engineer will ensure the availability, stability and reliability of AI GPU clusters, perform hardware and network troubleshooting, respond to production incidents, monitor cluster infrastructure health, and coordinate with OEM vendors to drive RMA cases through resolution.

Key Responsibilities

• Perform daily operations and maintenance of GPU/CPU servers, network switches, optical transceivers and related data center infrastructure, including troubleshooting, component replacement, rack-and-stack, cabling and configuration changes.

• Monitor GPU cluster health and infrastructure alerts, identify hardware and network abnormalities, perform initial diagnosis and remediation, and escalate issues to the appropriate internal teams or OEM vendors as required.

• Troubleshoot and resolve infrastructure incidents and complete post-incident reviews and Root Cause Analysis (RCA).

• Execute infrastructure changes according to established change-management procedures and maintain Standard Operating Procedures (SOPs), troubleshooting guides and operational documentation.

• Coordinate with hardware vendors for technical support, hardware replacement and RMA management, and track issues through final resolution.

Qualifications

• Degree or relevant educational background in Computer Science, Information Technology, Engineering or a related field.

• At least 2 years of hands-on experience in server, network or data center operations and maintenance.

• Familiarity with x86 server hardware and Linux operating systems, with a working understanding of TCP/IP and data center networking.

• Hands-on ability to diagnose hardware failures and replace components and other field-replaceable units.

• Hands-on experience with server BMC/OOB management interfaces such as iDRAC, IPMI or equivalent is preferred.

• Good communication skills, including the ability to read technical documentation and vendor knowledge bases, independently prepare RCA reports and emails, and communicate with OEM support teams through email, ticketing systems and phone calls.

• Strong sense of ownership, good communication skills and disciplined documentation practices.

• Experience with RDMA, InfiniBand, RoCE, high-performance networking or GPU clusters is strongly preferred.

• Experience with monitoring platforms such as Zabbix, Prometheus and Grafana is preferred.

Preferred Qualifications

• Hands-on operations experience with large-scale GPU clusters consisting of 1,000+ GPUs; experience supporting 10,000+ GPU environments is a strong plus.

• RHCE, CCNP or similar industry certifications.

• Troubleshooting experience with NVIDIA/Mellanox or Broadcom networking equipment.

• Experience working within ITIL-based incident, problem and change-management processes.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: PTPe6m9jp2JTIqN
  • Position Id: 9093066
  • Posted 5 days ago
Contact the job poster
SM

Sujuan Meng

Recruiter @ Aquila Hash, Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

New York

•

5d ago

Easy Apply

Full-time, Third Party

Depends on Experience

Georgia

•

Today

Full-time

USD 60,000.00 - 75,000.00 per year

Remote

•

Today

Full-time

Endicott, New York

•

Today

Full-time

Search all similar jobs