AI Infrastructure Engineer L3


HCLTech
Dice Job Match Score™
🔢 Crunching numbers...
Job Details
Skills
- Artificial Intelligence
- CUDA
- GPU
- Kubernetes
- Terraform
- AI Infrastructure
- NVIDIA GPU
- TensorRT
- HPC
- MLOps
- Kubeflow
- MLflow
- GPU Clusters
- AI Platform Engineering
- InfiniBand
- Ceph
- Deep Learning Infrastructure.
- "NVIDIA GPU"
- "GPU Engineer"
- "AI Infrastructure"
- "HPC"
- "Kubernetes"
- "CUDA"
- "TensorRT"
- "Kubeflow"
- "Ray"
- "MLflow"
- "InfiniBand"
- "RDMA"
- "NCCL"
- "DeepSpeed"
- "vLLM"
- "Triton Inference Server"
- "GPU Operator"
Summary
AI Infrastructure Engineer L3
Location: Santa Clara, CA
Experience: 10 plus Years
Employment Type: Full-Time
About HCLTech
HCLTech is a global technology company with over 220,000 professionals across 60 countries, delivering industry-leading capabilities in Digital, Engineering, Cloud, and AI. We help enterprises accelerate innovation through cutting-edge technologies and world-class talent.
Job Summary
We are seeking an experienced AI Infrastructure Engineer (L3) to design, deploy, optimize, and support high-performance AI and Machine Learning infrastructure. The ideal candidate will have deep expertise in GPU platforms, Kubernetes, HPC environments, distributed systems, and cloud-native AI technologies. This role involves managing large-scale GPU clusters, supporting AI training and inference workloads, troubleshooting complex infrastructure issues, and driving platform reliability.
Key Responsibilities
- Deploy and manage NVIDIA GPU infrastructure (A100, H100, L40) and AI accelerator platforms.
- Administer Kubernetes GPU clusters using NVIDIA GPU Operator and related technologies.
- Install and maintain CUDA, cuDNN, TensorRT, firmware, and driver stacks.
- Manage high-performance storage solutions such as Ceph, Lustre, BeeGFS, and NFS.
- Support InfiniBand, RDMA, RoCE, NVLink, and other high-speed networking technologies.
- Optimize Linux environments (RHEL, Ubuntu, Rocky Linux) for AI and HPC workloads.
- Support AI orchestration platforms including Kubeflow, MLflow, Ray, and Slurm.
- Implement Infrastructure as Code using Terraform, Helm, and GitOps tools.
- Monitor platform performance with Prometheus, Grafana, NVIDIA DCGM, and OpenTelemetry.
- Lead root cause analysis (RCA) and resolve critical GPU, networking, storage, and platform issues.
- Collaborate with cloud, data science, MLOps, SRE, and engineering teams to deliver scalable AI platforms.
Required Skills
- Strong experience with NVIDIA GPU platforms and GPU cluster administration.
- Expertise in Kubernetes, containerization, and cloud-native technologies.
- Hands-on experience with CUDA, TensorRT, NCCL, DeepSpeed, Horovod, and distributed training.
- Strong Linux administration and performance tuning skills.
- Experience with Terraform, Helm, ArgoCD, and automation frameworks.
- Knowledge of AI infrastructure, MLOps, and large-scale distributed systems.
- Excellent troubleshooting, debugging, and production support experience.
Preferred Certifications
- NVIDIA Certified Associate AI Infrastructure
- NVIDIA Base Command Manager Certification
- AWS Solutions Architect Associate
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Application Developer (CKAD)
Qualifications
- Bachelor's Degree in Computer Science, Engineering, or a related field.
- 8-12 years of Infrastructure or Platform Engineering experience.
- 4-6 years supporting AI/ML environments and GPU-based platforms.
- Experience operating production-scale AI infrastructure.
Disclaimer
HCL is an equal opportunity employer, committed to providing equal employment opportunities to all applicants and employees regardless of race, religion, sex, color, age, national origin, pregnancy, sexual orientation, physical disability or genetic information, military or veteran status, or any other protected classification, in accordance with federal, state, and/or local law. Should any applicant have concerns about discrimination in the hiring process, they should provide a detailed report of those concerns to for investigation.
Compensation and Benefits
A candidate s pay within the range will depend on their work location, skills, experience, education, and other factors permitted by law. This role may also be eligible for performance-based bonuses subject to company policies. In addition, this role is eligible for the following benefits subject to company policies: medical, dental, vision, pharmacy, life, accidental death & dismemberment, and disability insurance; employee assistance program; 401(k) retirement plan; 10 days of paid time off per year (some positions are eligible for need-based leave with no designated number of leave days per year); and 10 paid holidays per year.
- Dice Id: hcl001
- Position Id: 9068421
- Posted 22 hours ago
Company Info
HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering and cloud, powered by a broad portfolio of technology services and products.
We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending March 2025 totaled $13.8 billion.
To learn how we can supercharge progress for you, visit hcltech.com.

Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs