Job Title: AI Senior Engineer / Senior AI Infrastructure Engineer
Location: Plano, TX (hybrid)
Job Summary
We are seeking an experienced Senior AI Infrastructure Engineer to design, build, and optimize next-generation AI infrastructure, GPU cloud platforms, and accelerated computing environments. The ideal candidate will have strong expertise in NVIDIA AI infrastructure, AI Factory architectures, GPUaaS, and high-performance computing.
The engineer will work on large-scale AI data center solutions supporting Generative AI, Large Language Models (LLMs), AI cloud platforms, and enterprise AI workloads. The role requires hands-on experience with NVIDIA reference architectures, DGX, HGX, Blackwell, GB200, Spectrum-X, and InfiniBand technologies.
Key Responsibilities
- Design and architect scalable AI infrastructure for enterprise AI, GenAI, and LLM workloads.
- Develop and implement AI Factory and GPUaaS (GPU as a Service) platforms.
- Design GPU-accelerated computing environments using NVIDIA reference architectures.
- Work with NVIDIA DGX, HGX, Blackwell GPU platforms, and GB200 systems.
- Architect high-performance GPU clusters and AI data center infrastructure.
- Design and optimize high-speed networking using NVIDIA Spectrum-X and InfiniBand.
- Support AI cloud and GPU cloud infrastructure deployment, integration, and operations.
- Collaborate with data center, cloud, networking, storage, and platform engineering teams.
- Evaluate infrastructure requirements for AI training, inference, and large-scale model deployment.
- Develop infrastructure standards, technical designs, and deployment documentation.
- Troubleshoot performance, scalability, networking, and infrastructure integration issues.
- Support capacity planning, resource utilization, and optimization of GPU infrastructure.
- Contribute to the design and implementation of LLM infrastructure and GenAI platforms.
- Work with engineering and architecture teams to evaluate emerging NVIDIA AI technologies.
Required Technical Skills
- Strong experience in AI Infrastructure and Accelerated Computing.
- Experience with AI Factory architecture and GPUaaS platforms.
- Strong knowledge of NVIDIA Reference Architecture.
- Hands-on experience with NVIDIA DGX and/or HGX systems.
- Experience with NVIDIA Blackwell architecture and/or GB200 platforms.
- Knowledge of NVIDIA Spectrum-X networking.
- Strong experience with InfiniBand networking and high-performance GPU clusters.
- Experience designing or supporting AI cloud, GPU cloud, or AI data center environments.
- Understanding of LLM infrastructure, Generative AI platforms, and AI workload deployment.
- Experience with infrastructure architecture, performance optimization, and system integration.
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
- Experience with NVIDIA AI Enterprise, CUDA, or GPU software ecosystems.
- Experience with high-performance computing (HPC) and distributed AI workloads.
- Knowledge of GPU cluster management, containerized AI workloads, and cloud infrastructure.
- Experience with large-scale AI data center deployments.
- Strong analytical, problem-solving, and communication skills.