Senior LLMOps / MLOps Engineer

Santa Clara, CA, US • Posted 2 hours ago • Updated 2 hours ago
Contract W2
Contract Corp To Corp
6 Months
Travel Required
On-site
$55 - $65/hr
Fitment

Dice Job Match Score™

⭐ Evaluating experience...

Job Details

Skills

  • MLOps
  • LLMOps
  • AI Platform
  • Generative AI
  • Model Serving
  • GPU Optimization
  • Python
  • AI/ML Platform Engineering
  • LLM
  • vLLM
  • SGLang
  • TGI
  • Triton
  • Ray Serve
  • Azure ML
  • Databricks Model Serving

Summary

Role: Senior LLMOps / MLOps Engineer Location: Santa Clara, CA (Onsite) Duration: 6-12+ Months Contract

Only Local Candidates will work out

Must Have Skills:
Skill 1 Strong proficiency in Python and software engineering best practices
Skill 2 14+ years of experience in MLOps, LLMOps, AI/ML Platform Engineering
Skill 3 Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving

Good To have Skills:
Skill 1 Exposure to AI Observability, Governance, and Responsible AI practices

Summary:
We are looking for a highly skilled Senior LLMOps / MLOps Engineer with strong expertise in LLM inferencing, model hosting, and serving Large Language Models (LLMs) at scale. The ideal candidate should be a hands-on engineer with proven experience deploying and optimizing open-source LLMs, building high-performance inference platforms using technologies such as vLLM, SGLang, TGI, Triton, and Ray Serve, and driving GPU utilization, latency, throughput, and cost optimization. This is a highly technical role requiring active involvement in designing, building, troubleshooting, and optimizing production AI systems. Experience in MLOps platforms and scalable AI infrastructure is essential.

Must-Have Skills:
5-7 years of experience in MLOps, LLMOps, AI/ML Platform Engineering.
Strong proficiency in Python and software engineering best practices.
Experience working with open-source LLMs such as Llama, Mistral, Gemma, or Qwen.
Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving.
Experience with Kubernetes, Docker, Azure ML, Databricks, and MLflow.
Good understanding of RAG, Vector Databases, GPU Optimization, Quantization, KV Cache, PagedAttention, and ContinuoDynamic Batching.
Demonstrated hands-on experience building, deploying, troubleshooting, and optimizing production-grade LLM and GenAI solutions.
Experience deploying, scaling, and monitoring production-grade GenAI/LLM applications.
Exposure to AI Observability, Governance, and Responsible AI practices.

Good-to-Have Skills:
Hands-on experience with LLM Fine-Tuning using PEFT, SFT, CPT, LoRA, and QLoRA techniques.
Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT.
Knowledge of distributed training and multi-GPU environments.
Experience with Agentic AI frameworks such as LangGraph, AutoGen, or CrewAI.
Understanding of simulation platforms, digital twins, modeling & simulation workflows, or scientific computing.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10121181
  • Position Id: 9044699
  • Posted 2 hours ago
Contact the job poster
AM

Abhishek Mishra

Recruiter @ Cardinal Integrated Technologies Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Santa Clara, California

Today

Easy Apply

Contract, Third Party

Santa Clara, California

5d ago

Easy Apply

Third Party, Contract

Depends on Experience

Santa Clara, California

Today

Full-time

USD 152,925.00 - 254,875.00 per year

Santa Clara, California

Today

Full-time

USD 178,500.00 per year

Search all similar jobs