AI Cluster Validation Engineer (Tester & Programmer)

Morrisville, NC, US • Posted 1 hour ago • Updated 1 hour ago
Full Time
On-site
$100,000 - $120,000/yr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • AI/HPC GPU

Summary

Key Responsibilities

  • AI Cluster Infrastructure Validation: Build, deploy, and maintain hardware and software test environments for AI cluster validation, qualification, and performance benchmarking.
  • Execute test plans, reproduce complex issues, collect and analyze logs, and perform first-level root cause analysis across hardware, software, networking, and system components.
  • Develop and enhance automated test frameworks, scripts, and tools to improve validation efficiency, coverage, and repeatability.
  • Collaborate with software, hardware, networking, and system engineering teams to investigate issues, validate fixes, and improve overall cluster stability, scalability, and performance.
  • Conduct functional, performance, stress, and reliability testing for AI/HPC cluster solutions.
  • Document test methodologies, configurations, results, troubleshooting procedures, and operational best practices.

Basic Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
  • 1+ years of experience in deploying, testing, validating, or supporting data center hardware and software systems, with expertise in one or more of the following areas:

Server systems, Networking, Storage

  • Strong understanding of Linux operating systems, system administration, and troubleshooting.
  • Familiarity with data center cluster management platforms, distributed computing environments, and infrastructure validation methodologies.
  • Programming or scripting experience with Python, Bash, or similar languages.
  • Strong analytical, debugging, problem-solving, and troubleshooting skills.
  • Proven ability to quickly learn new technologies and adapt in a fast-paced engineering environment.
  • Self-motivated team player with strong communication and collaboration skills.

Preferred Qualifications

  • Experience operating, maintaining, or supporting data center, cloud, or laboratory environments.
  • Familiarity with AI/HPC clusters, GPU-based systems, and high-speed interconnect technologies such as InfiniBand, RoCE, NVLink, and Ethernet fabrics.
  • Experience developing test automation tools, validation frameworks, or CI/CD pipelines.
  • Knowledge of cluster orchestration and management technologies, including Kubernetes, Slurm, virtualization platforms, or cloud infrastructure.
  • Experience with performance analysis, benchmarking, workload characterization, and system optimization.
  • Understanding of storage technologies, distributed file systems, and AI workload deployment environments.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 501494924
  • Position Id: 9084003
  • Posted 1 hour ago
Contact the job poster
Paras Premchand Lallwani

Paras Premchand Lallwani

Recruitments | North America Talent Acquisition Group at Mphasis - USA Hiring @ MphasiS Corporation USA
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Morrisville, North Carolina

Today

Easy Apply

Full-time

$100,000 - $120,000

Georgia

Today

Full-time

USD 60,000.00 - 75,000.00 per year

Remote

Today

Easy Apply

Contract

Up to $90

Georgia

Today

Full-time

USD 76,000.00 - 94,000.00 per year

Search all similar jobs