Job Title: AI Data Center Architect
Location: Plano, TX - Hybrid
Experience: 10+ years in Data Center, Cloud Infrastructure, HPC, or AI Infrastructure Architecture
Job Summary
We are seeking an experienced AI Data Center Architect to design, architect, and lead the implementation of next-generation AI infrastructure and AI Factory environments. The ideal candidate will have strong expertise in GPU-based computing, NVIDIA Reference Architecture, GPUaaS, accelerated computing, AI cloud platforms, and large-scale data center infrastructure.
The architect will be responsible for developing scalable, high-performance, and reliable infrastructure for Generative AI, Large Language Models (LLM), machine learning, and enterprise AI workloads.
The role requires hands-on architectural experience with NVIDIA DGX, HGX, Blackwell, GB200, Spectrum-X, InfiniBand, and GPU cloud infrastructure.
Key Responsibilities
AI Infrastructure & Data Center Architecture
- Design end-to-end AI Data Center and AI Factory architectures for large-scale GPU workloads.
- Develop infrastructure blueprints for AI training, inference, and accelerated computing environments.
- Architect scalable GPU clusters, compute, networking, storage, and data center infrastructure.
- Evaluate infrastructure requirements for LLM, GenAI, and enterprise AI platforms.
- Design high-availability, scalable, and performance-optimized AI infrastructure solutions.
- Define infrastructure standards, reference designs, and architecture best practices.
NVIDIA AI Infrastructure
- Design and implement NVIDIA-based AI infrastructure using NVIDIA Reference Architecture.
- Architect NVIDIA DGX and HGX platforms for enterprise AI and large-scale GPU computing.
- Experience with NVIDIA Blackwell architecture and GB200 systems.
- Design GPU clusters and accelerated computing platforms for AI workloads.
- Evaluate NVIDIA GPU, CPU, networking, storage, and software stack requirements.
- Collaborate with hardware, cloud, network, and data center engineering teams.
GPUaaS & AI Cloud
- Architect GPU-as-a-Service (GPUaaS) and GPU Cloud platforms.
- Design multi-tenant GPU infrastructure and resource allocation models.
- Develop scalable AI Cloud and GenAI platform architectures.
- Support GPU provisioning, scheduling, monitoring, and capacity planning.
- Define infrastructure requirements for AI model training, inference, and deployment.
- Evaluate cloud and on-premises GPU infrastructure solutions.
High-Performance Networking
- Architect high-performance networking for AI GPU clusters.
- Strong experience with NVIDIA Spectrum-X networking and InfiniBand.
- Design low-latency, high-bandwidth network architectures for distributed AI workloads.
- Evaluate RDMA, RoCE, InfiniBand, Ethernet, and GPU interconnect technologies.
- Support network scalability, performance optimization, and troubleshooting.
- Collaborate with network engineering teams on AI data center deployments.
AI Platform & LLM Infrastructure
- Design infrastructure platforms supporting Generative AI and LLM workloads.
- Define compute, GPU, storage, networking, and orchestration requirements.
- Support AI model training, fine-tuning, inference, and serving environments.
- Collaborate with AI/ML engineers, platform engineers, DevOps, and cloud architects.
- Evaluate containerized AI infrastructure, Kubernetes, and GPU orchestration platforms.
- Develop infrastructure strategies for enterprise AI and AI cloud environments.
Architecture Governance & Delivery
- Create High-Level Design (HLD), Low-Level Design (LLD), architecture diagrams, and technical documentation.
- Lead architecture reviews, technology evaluations, and infrastructure design workshops.
- Define capacity planning, performance, availability, scalability, and security requirements.
- Support proof of concepts, infrastructure deployment, and production readiness.
- Work with vendors and technology partners on infrastructure solutions.
- Provide technical leadership and guidance to infrastructure engineering teams.
Required Technical Skills
- 10+ years of experience in Data Center Architecture, Cloud Infrastructure, HPC, or AI Infrastructure.
- Strong experience designing AI Factory or AI Data Center environments.
- Expertise in NVIDIA Reference Architecture.
- Hands-on architecture experience with NVIDIA DGX and HGX platforms.
- Experience with NVIDIA Blackwell architecture and GB200 systems.
- Strong knowledge of GPUaaS, GPU Cloud, and AI Cloud infrastructure.
- Expertise in AI Infrastructure and Accelerated Computing.
- Experience with NVIDIA Spectrum-X and InfiniBand networking.
- Strong understanding of LLM Infrastructure and GenAI Platform architecture.
- Experience designing scalable GPU clusters and distributed AI environments.
- Strong knowledge of data center compute, networking, storage, and infrastructure architecture.
Preferred Qualifications
- Experience with NVIDIA AI Enterprise and AI software ecosystems.
- Experience with Kubernetes, GPU orchestration, and containerized AI platforms.
- Knowledge of RDMA, RoCE, NVLink, and high-performance GPU interconnects.
- Experience with AI infrastructure monitoring, automation, and observability.
- Experience with enterprise AI cloud or hyperscale data center environments.
- Knowledge of power, cooling, rack density, and physical data center requirements.
- NVIDIA certifications or relevant cloud/data center architecture certifications.
- Bachelor’s or Master’s degree in Computer Science, Engineering, Information Technology, or a related field.