Senior LLMOps / MLOps Engineer :: Santa Clara, CA-5Day onsite. -Only locals.


American IT Systems
Dice Job Match Score™
🫥 Flibbertigibetting...
Job Details
Skills
- building
- troubleshooting
- We are looking for a highly skilled Senior LLMOps / MLOps Engineer with strong expertise in LLM inferencing
- model hosting
- and serving Large Language Models (LLMs) at scale. The ideal candidate should be a hands-on engineer with proven experience deploying and optimizing open-source LLMs
- building high-performance inference platforms using technologies such as vLLM
- SGLang
- TGI
- Triton
- and Ray Serve
- and driving GPU utilization
- latency
- throughput
- and cost optimization. This is a highly technical role requiring active involvement in designing
- and optimizing production AI systems. Experience in MLOps platforms and scalable AI infrastructure is essential.
Summary
Try to submit locals within 50 miles -3 resumes
Senior LLMOps / MLOps Engineer
Location: Santa Clara, CA 5days onsite
Summary
We are looking for a highly skilled Senior LLMOps / MLOps Engineer with strong expertise in LLM inferencing, model hosting, and serving Large Language Models (LLMs) at scale.
The ideal candidate should be a hands-on engineer with proven experience deploying and optimizing open-source LLMs, building high-performance inference platforms using technologies such as vLLM, SGLang, TGI, Triton, and Ray Serve, and driving GPU utilization, latency, throughput, and cost optimization.
This is a highly technical role requiring active involvement in designing, building, troubleshooting, and optimizing production AI systems.
Experience in MLOps platforms and scalable AI infrastructure is essential.
Must-Have Skills
5-7 years of experience in MLOps, LLMOps, AI/ML Platform Engineering.
Strong proficiency in Python and software engineering best practices.
Experience working with open-source LLMs such as Llama, Mistral, Gemma, or Qwen.
Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving.
Experience with Kubernetes, Docker, Azure ML, Databricks, and MLflow.
Good understanding of RAG, Vector Databases, GPU Optimization, Quantization, KV Cache, PagedAttention, and ContinuoDynamic Batching.
Demonstrated hands-on experience building, deploying, troubleshooting, and optimizing production-grade LLM and GenAI solutions.
Experience deploying, scaling, and monitoring production-grade GenAI/LLM applications.
Exposure to AI Observability, Governance, and Responsible AI practices.
Good-to-Have Skills
Hands-on experience with LLM Fine-Tuning using PEFT, SFT, CPT, LoRA, and QLoRA techniques.
Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT.
Knowledge of distributed training and multi-GPU environments.
Experience with Agentic AI frameworks such as LangGraph, AutoGen, or CrewAI.
Understanding of simulation platforms, digital twins, modeling & simulation workflows, or scientific computing.
- Dice Id: 91163020
- Position Id: 2026-1444
- Posted 10 hours ago
Company Info
About American IT Systems
American IT Systems Staffing Fastest growing Recruitment firm helping clients hire the best quality candidates faster across the globe.
We provide services starting from Temporary Staffing & Permanent Recruitment to management consulting. We specialize in Technology, Product & Design hiring headhunting & sourcing passive resources using cutting edge technology & tools
Our Aim
Hiring and recruiting top talent can be a challenging, time-intensive process. Client organizations recognize the value in saving time, money and preventing unnecessary burden on internal staff by outsourcing certain hiring needs. It can also send a strong message to top professionals they will spare no expense in finding and hiring the best talent possible..


Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs