Maxonic maintains a close and long-term relationship with our direct client. In support of their needs, we are looking for an AI Devops Infrastructure Engineer/GPU Infrastructure Engineer.
Job Title: AI Devops Infrastructure Engineer/GPU Infrastructure Engineer
Job Type: Contract / Contract to Hire
Job Location: San Jose, CA
Work Schedule: Onsite
Pay Rate: $100 on W2 & $100-124 on c2c
Description:
We are looking for an AI Infrastructure Engineer to help build and operationalize infrastructure supporting AI development, experimentation, training, inference, and AI-enabled developer productivity. This role will focus on building scalable, highly resilient GPU-based infrastructure and enabling engineering teams to access AI resources through reliable, secure, and automated infrastructure services.
The ideal candidate has hands-on experience building and operating AI/ML infrastructure in real-world environments, with a strong foundation in infrastructure engineering, DevOps/SRE, platform engineering, or similar disciplines.
Responsibilities
- Build and operate AI infrastructure supporting development, experimentation, training, and inference.
- Design and manage infrastructure across GPU and accelerator environments, including NVIDIA and/or AMD.
- Build scalable infrastructure for GPU capacity planning, utilization, forecasting, workload scheduling, resource allocation, and performance optimization.
- Develop self-service AI infrastructure capabilities that allow engineering teams to request, provision, manage, extend, and release GPU, CPU, memory, and storage resources.
- Build and maintain infrastructure across Kubernetes and containerized environments, including compute, networking, storage, and accelerator scheduling.
- Automate infrastructure provisioning, configuration, scaling, patching, driver/firmware lifecycle management, and decommissioning.
- Use Infrastructure-as-Code, APIs, automation, and scripting to improve infrastructure reliability and operational efficiency.
- Establish monitoring, observability, alerting, capacity management, and operational processes for AI infrastructure.
- Help define and implement best practices for scalable, feasible, and operationally sustainable AI infrastructure.
- Integrate AI infrastructure with Developer Productivity tooling, CI/CD pipelines, source control, artifact management, build infrastructure, and developer tooling.
- Partner with engineering, security, IT, and AI/ML teams to make AI resources accessible with minimal operational friction.
- Troubleshoot complex infrastructure issues and perform root-cause analysis.
- Help establish secure infrastructure standards covering network segmentation, access controls, authentication, authorization, secrets management, and governance.
- Evaluate emerging AI infrastructure technologies and improve scalability, reliability, performance, automation, and developer experience.
Qualifications:
- Strong experience in Infrastructure Engineering, Platform Engineering, DevOps, SRE, Cloud Engineering, or AI/ML Infrastructure.
- Hands-on experience building and operating AI infrastructure, not just exposure to AI/ML concepts.
- Strong experience with GPU infrastructure and GPU scalability.
- Experience with Kubernetes and containerized production environments.
- Strong understanding of compute, networking, storage, workload scheduling, and resource allocation.
- Experience with GPU capacity planning, utilization, forecasting, workload scheduling, and performance optimization.
- Experience building or operating self-service infrastructure/provisioning platforms.
- Strong experience with Infrastructure-as-Code and automation, such as Terraform, Ansible, Python, APIs, or similar technologies.
- Experience with CI/CD pipelines and infrastructure operations.
- Experience with monitoring and observability tools such as Grafana, Prometheus, OpenTelemetry, Kibana, or similar platforms.
- Experience working across bare-metal, private cloud, and/or public cloud environments.
- Familiarity with AI/ML infrastructure, inference platforms, and distributed AI workloads.
- Strong troubleshooting, root-cause analysis, communication, and cross-functional collaboration skills.
About Maxonic:
Since 2002 Maxonic has been at the forefront of connecting candidate strengths to client challenges. Our award winning, dedicated team of recruiting professionals are specialized by technology, are great listeners, and will seek to find a position that meets the long-term career needs of our candidates. We take pride in the over 10,000 candidates that we have placed, and the repeat business that we earn from our satisfied clients.
Interested in Applying?
Please apply with your most current resume. Feel free to contact Pramod Kumar (/) for more details.