Job title: Senior AI/ML Ops Engineer
Work location – Cupertino, CA-Hybrid
Contract duration -Long term contract
Visa:- / only
Skill Set
Job Description
Senior AI/ML Ops Engineer
Agent Evaluation, Observability & Production Reliability
About the Role
We are looking for a Senior AI/ML Ops Engineer to own the evaluation and reliability backbone of our AI agent ecosystem. You will design and build a comprehensive evaluation framework from the ground up, then use it to measure, monitor, and continuously improve a portfolio of production LLM-powered agents. This is a high-visibility role that sits at the intersection of ML engineering, platform operations, and quality: when a critical agent misbehaves in production, you will be the person teams look to for fast, rigorous triage and durable resolution.
You will work cross-functionally with agent developers, product owners, and infrastructure teams to define what “good” looks like for each agent, encode it into automated evaluations, wire those evaluations into CI/CD so regressions are caught before they ship, and build the reporting that gives leadership a clear, trustworthy view of agent quality over time.
What You’ll Do
- Design and build an end-to-end evaluation framework for LLM-based agents, covering offline evaluation, regression testing, online monitoring, and human-in-the-loop review.
- Develop and maintain evaluation suites across multiple agents: golden datasets, task-completion and accuracy metrics, LLM-as-judge rubrics, safety and guardrail checks, latency and cost benchmarks.
- Integrate evaluations into CI/CD pipelines as quality gates, so every prompt, model, tool, or orchestration change is automatically evaluated before promotion to production.
- Build monitoring, alerting, and observability for agents in production, including tracing of multi-step agent runs, drift detection, and anomaly detection on quality and behavioral metrics.
- Create dashboards and recurring reporting that communicate agent performance, quality trends, incident history, and evaluation coverage to engineering teams and leadership.
- Lead the triage and resolution of highly complex, critical agent issues in production: reproduce failures, isolate root causes across prompts, models, tools, retrieval, and infrastructure, and drive fixes through to verified resolution.
- Partner cross-functionally with agent developers, product, and platform teams to define acceptance criteria, prioritize fixes, and establish runbooks, severity levels, and escalation paths for agent incidents.
- Establish and champion best practices for agent versioning, release management, rollback, canary deployments, and A/B evaluation of agent changes.
- Continuously improve the evaluation platform itself: expand coverage, reduce evaluation runtime and cost, and automate away manual review wherever quality allows.
Must Have
- 5+ years of experience in ML engineering, MLOps, platform engineering, or SRE, including 2+ years working hands-on with LLMs or LLM-powered applications in production.
- Demonstrated experience building evaluation systems for ML or LLM applications: test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.
- Strong software engineering skills in Python (and ideally TypeScript), with a track record of building reliable, well-tested internal platforms and tooling.
- Deep familiarity with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite) and experience embedding automated quality gates into deployment pipelines.
- Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus) and, ideally, LLM-specific observability tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave).
- Proven ability to debug complex distributed systems under pressure, including production incident response, root-cause analysis, and blameless postmortems.
- Excellent cross-functional communication: able to translate evaluation results into clear findings and recommendations for both engineers and non-technical stakeholders.
- Comfort with ambiguity and a builder’s mindset: this role starts with a blank page and ends with the evaluation platform the whole organization relies on.
- Experience with agentic frameworks and orchestration patterns (e.g., multi-agent systems, tool use, RAG pipelines) and their distinct failure modes.
- Experience with prompt management, model routing, or fine-tuning workflows and evaluating changes across model versions and providers.
- Background in statistics or experimentation (A/B testing, significance testing, sampling strategies for human review).
- Design and build reusable AI agent skills, plugins and maintain internal marketplace infrastructure to extend and scale Data, AIML capabilities across the organization.
- Expertise in causal inference and measurement strategy including causal graphs, ontologies, and knowledge graphs to drive rigorous, decision grade data analysis.
- Experience operating in regulated or high-stakes domains where agent errors carry real business or customer impact.
- Success in the first 90 days
- A production-grade evaluation framework is in place, with automated evaluation suites covering every critical agent.
- CI/CD quality gates catch agent regressions before release, with clear pass/fail criteria trusted by agent teams.
- Production agents have real-time quality monitoring and alerting, with documented runbooks and severity-based escalation paths.
- Leadership receives regular, reliable reporting on agent quality, and time-to-resolution for critical agent incidents has measurably decreased.