Senior AI/ML Ops Engineer

Hybrid in Cupertino, CA, US • Posted 1 day ago • Updated 1 day ago
Contract W2
Contract Independent
12 Months
No Travel Required
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • ML
  • MLOps
  • platform engineering

Summary

Job title: Senior AI/ML Ops Engineer

Work location – Cupertino, CA-Hybrid

Contract duration -Long term contract

Visa:- / only

 

Skill Set

Job Description

Senior AI/ML Ops Engineer

Agent Evaluation, Observability & Production Reliability

 

About the Role

We are looking for a Senior AI/ML Ops Engineer to own the evaluation and reliability backbone of our AI agent ecosystem. You will design and build a comprehensive evaluation framework from the ground up, then use it to measure, monitor, and continuously improve a portfolio of production LLM-powered agents. This is a high-visibility role that sits at the intersection of ML engineering, platform operations, and quality: when a critical agent misbehaves in production, you will be the person teams look to for fast, rigorous triage and durable resolution.

You will work cross-functionally with agent developers, product owners, and infrastructure teams to define what “good” looks like for each agent, encode it into automated evaluations, wire those evaluations into CI/CD so regressions are caught before they ship, and build the reporting that gives leadership a clear, trustworthy view of agent quality over time.

 

What You’ll Do

  • Design and build an end-to-end evaluation framework for LLM-based agents, covering offline evaluation, regression testing, online monitoring, and human-in-the-loop review.
  • Develop and maintain evaluation suites across multiple agents: golden datasets, task-completion and accuracy metrics, LLM-as-judge rubrics, safety and guardrail checks, latency and cost benchmarks.
  • Integrate evaluations into CI/CD pipelines as quality gates, so every prompt, model, tool, or orchestration change is automatically evaluated before promotion to production.
  • Build monitoring, alerting, and observability for agents in production, including tracing of multi-step agent runs, drift detection, and anomaly detection on quality and behavioral metrics.
  • Create dashboards and recurring reporting that communicate agent performance, quality trends, incident history, and evaluation coverage to engineering teams and leadership.
  • Lead the triage and resolution of highly complex, critical agent issues in production: reproduce failures, isolate root causes across prompts, models, tools, retrieval, and infrastructure, and drive fixes through to verified resolution.
  • Partner cross-functionally with agent developers, product, and platform teams to define acceptance criteria, prioritize fixes, and establish runbooks, severity levels, and escalation paths for agent incidents.
  • Establish and champion best practices for agent versioning, release management, rollback, canary deployments, and A/B evaluation of agent changes.
  • Continuously improve the evaluation platform itself: expand coverage, reduce evaluation runtime and cost, and automate away manual review wherever quality allows.

 

Must Have

  • 5+ years of experience in ML engineering, MLOps, platform engineering, or SRE, including 2+ years working hands-on with LLMs or LLM-powered applications in production.
  • Demonstrated experience building evaluation systems for ML or LLM applications: test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.
  • Strong software engineering skills in Python (and ideally TypeScript), with a track record of building reliable, well-tested internal platforms and tooling.
  • Deep familiarity with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite) and experience embedding automated quality gates into deployment pipelines.
  • Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus) and, ideally, LLM-specific observability tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave).
  • Proven ability to debug complex distributed systems under pressure, including production incident response, root-cause analysis, and blameless postmortems.
  • Excellent cross-functional communication: able to translate evaluation results into clear findings and recommendations for both engineers and non-technical stakeholders.
  • Comfort with ambiguity and a builder’s mindset: this role starts with a blank page and ends with the evaluation platform the whole organization relies on.
  • Experience with agentic frameworks and orchestration patterns (e.g., multi-agent systems, tool use, RAG pipelines) and their distinct failure modes.
  • Experience with prompt management, model routing, or fine-tuning workflows and evaluating changes across model versions and providers.
  • Background in statistics or experimentation (A/B testing, significance testing, sampling strategies for human review).
  • Design and build reusable AI agent skills, plugins and maintain internal marketplace infrastructure to extend and scale Data, AIML capabilities across the organization.
  • Expertise in causal inference and measurement strategy including causal graphs, ontologies, and knowledge graphs to drive rigorous, decision grade data analysis.
  • Experience operating in regulated or high-stakes domains where agent errors carry real business or customer impact.
  • Success in the first 90 days
  • A production-grade evaluation framework is in place, with automated evaluation suites covering every critical agent.
  • CI/CD quality gates catch agent regressions before release, with clear pass/fail criteria trusted by agent teams.
  • Production agents have real-time quality monitoring and alerting, with documented runbooks and severity-based escalation paths.
  • Leadership receives regular, reliable reporting on agent quality, and time-to-resolution for critical agent incidents has measurably decreased.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: mategr
  • Position Id: Nilesh_Payal95
  • Posted 1 day ago
Contact the job poster
Murtuza Tajir

Murtuza Tajir

Recruiter @ ClifyX
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Cupertino, California

Today

Easy Apply

Contract

Depends on Experience

Cupertino, California

Today

Easy Apply

Third Party, Contract

$45 - $50

Mountain View, California

Today

Full-time

USD 193,930.00 - 352,290.00 per year

Palo Alto, California

11d ago

Easy Apply

Contract, Third Party

$80 - $90

Search all similar jobs