Senior AI/ML Ops Engineer Agent Evaluation, Observability & Production Reliability

Austin, TX, US • Posted 8 hours ago • Updated 8 hours ago
Contract W2
Contract Independent
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Artificial Intelligence
  • A/B Testing
  • Continuous Delivery
  • Continuous Integration
  • GitHub
  • GitLab
  • Machine Learning Operations (ML Ops)
  • Machine Learning (ML)
  • LangSmith
  • Grafana
  • Prometheus
  • Python
  • Stacks Blockchain
  • Root Cause Analysis
  • LLMs

Summary

Job Title: Senior AI/ML Ops Engineer Agent Evaluation, Observability & Production Reliability
Location: Cupertino, CA or Austin, TX (onsite)
Type: Contract Position

Job Description

Must Have

  • 5+ years of experience in ML engineering, MLOps, platform engineering, or SRE, including 2+ years working hands-on with LLMs or LLM-powered applications in production.
  • Demonstrated experience building evaluation systems for ML or LLM applications: test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.
  • Strong software engineering skills in Python (and ideally TypeScript), with a track record of building reliable, well-tested internal platforms and tooling.
  • Deep familiarity with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite) and experience embedding automated quality gates into deployment pipelines.
  • Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus) and, ideally, LLM-specific observability tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave).
  • Proven ability to debug complex distributed systems under pressure, including production incident response, root-cause analysis, and blameless postmortems.
  • Excellent cross-functional communication: able to translate evaluation results into clear findings and recommendations for both engineers and non-technical stakeholders.
  • Comfort with ambiguity and a builder s mindset: this role starts with a blank page and ends with the evaluation platform the whole organization relies on.
  • Experience with agentic frameworks and orchestration patterns (e.g., multi-agent systems, tool use, RAG pipelines) and their distinct failure modes.
  • Experience with prompt management, model routing, or fine-tuning workflows and evaluating changes across model versions and providers.
  • Background in statistics or experimentation (A/B testing, significance testing, sampling strategies for human review).
  • Design and build reusable AI agent skills, plugins and maintain internal marketplace infrastructure to extend and scale Data, AIML capabilities across the organization.
  • Expertise in causal inference and measurement strategy including causal graphs, ontologies, and knowledge graphs to drive rigorous, decision grade data analysis.
  • Experience operating in regulated or high-stakes domains where agent errors carry real business or customer impact.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91137892
  • Position Id: 9060421
  • Posted 8 hours ago
Contact the job poster
RC

Rashmi Chandak

Recruiter @ iPeople Infosystems LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Austin, Texas

Today

Easy Apply

Contract

Depends on Experience

Hybrid in Austin, Texas

Today

Easy Apply

Contract

Depends on Experience

Austin, Texas

Today

Full-time

USD 224,000.00 - 279,000.00 per year

Austin, Texas

Today

Full-time

USD 149,200.00 - 251,576.00 per year

Search all similar jobs