Position Overview
Responsible for the operations and maintenance of large-scale AI GPU computing clusters, including NVIDIA, AMD and/or other GPU server platforms, as well as associated Data Center Network (DCN), Out-of-Band Management (OOB) network, and high-performance GPU fabric, including InfiniBand and RDMA Ethernet.
The engineer will ensure the availability, stability and reliability of AI GPU clusters, perform hardware and network troubleshooting, respond to production incidents, monitor cluster infrastructure health, and coordinate with OEM vendors to drive RMA cases through resolution.
Key Responsibilities
• Perform daily operations and maintenance of GPU/CPU servers, network switches, optical transceivers and related data center infrastructure, including troubleshooting, component replacement, rack-and-stack, cabling and configuration changes.
• Monitor GPU cluster health and infrastructure alerts, identify hardware and network abnormalities, perform initial diagnosis and remediation, and escalate issues to the appropriate internal teams or OEM vendors as required.
• Troubleshoot and resolve infrastructure incidents and complete post-incident reviews and Root Cause Analysis (RCA).
• Execute infrastructure changes according to established change-management procedures and maintain Standard Operating Procedures (SOPs), troubleshooting guides and operational documentation.
• Coordinate with hardware vendors for technical support, hardware replacement and RMA management, and track issues through final resolution.
Qualifications
• Degree or relevant educational background in Computer Science, Information Technology, Engineering or a related field.
• At least 2 years of hands-on experience in server, network or data center operations and maintenance.
• Familiarity with x86 server hardware and Linux operating systems, with a working understanding of TCP/IP and data center networking.
• Hands-on ability to diagnose hardware failures and replace components and other field-replaceable units.
• Hands-on experience with server BMC/OOB management interfaces such as iDRAC, IPMI or equivalent is preferred.
• Good communication skills, including the ability to read technical documentation and vendor knowledge bases, independently prepare RCA reports and emails, and communicate with OEM support teams through email, ticketing systems and phone calls.
• Strong sense of ownership, good communication skills and disciplined documentation practices.
• Experience with RDMA, InfiniBand, RoCE, high-performance networking or GPU clusters is strongly preferred.
• Experience with monitoring platforms such as Zabbix, Prometheus and Grafana is preferred.
Preferred Qualifications
• Hands-on operations experience with large-scale GPU clusters consisting of 1,000+ GPUs; experience supporting 10,000+ GPU environments is a strong plus.
• RHCE, CCNP or similar industry certifications.
• Troubleshooting experience with NVIDIA/Mellanox or Broadcom networking equipment.
• Experience working within ITIL-based incident, problem and change-management processes.