Senior LLMOps / AgentOps Engineer Santa Clara - CA - California

Santa Clara, CA, US • Posted 1 day ago • Updated 1 day ago
Contract Corp To Corp
Contract W2
12 Months
On-site
$49 - $49/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • LLMOps
  • Google Cloud
  • Agent Engine
  • Model Armor

Summary

Role Summary

Senior subject-matter expert for the operations, observability, and lifecycle management of AI agents in production ( AgentOps LLMOps). Owns the frameworks and practices to safely deploy, monitor, evaluate, and continuously improve live agents ensuring reliability, safety, cost-efficiency, and business-KPI performance across the Intel Agent Factory.

Key Responsibilities

Define and operate the AgenticOps framework: agent registry, versioning, guarded rollout, and rollback for production agents.

Establish continuous evaluation and monitoring: quality, autonomy, safety (guardrails, Model Armor), latency, cost, and reuse metrics.

Implement observability and tracing for multi-agent systems (Agent Engine Observability, Cloud MonitoringLoggingTrace).

Own the 5-gate validation-to-production process and post-release escape management for delivered agents.

Design human-in-the-loop (HITL) supervision, feedback loops, and automated pre-production simulations for safe rollout.

Track and report agent business KPIs (CSAT, TAT, MTTR, cost savings) via AgentScore Agent 360 dashboards.

Drive cost governance for agent runtimes: model tiering, context caching, batchflex inference, budget caps and alerts.

Collaborate with DevOps SME (deploy) and AI & Data SME (grounding) to close the build-deploy-operate-improve loop advise Intel on AgenticOps ownership transfer.

Mandatory (Must-Have) Skills

Strong LLMOps MLOps AgentOps experience operating GenAI or agentic systems in production.

Hands-on with Google Cloud agent runtimes: Vertex AI, Agent Engine, and observability tooling.

Agent evaluation and safety: eval frameworks, guardrails, Model Armor, HITL, promptrobustness testing.

Monitoring, tracing, and reliability engineering (SRE) for AI workloads.

Cost governance and performance tuning for LLMagent workloads.

Proficiency in Python strong grasp of agent lifecycle and governance.

Preferred (Good-to-Have) Skills

Experience with ADK, A2A, MCP, and multi-agent orchestration in production.

BigQueryLooker for agent analytics and KPI dashboards.

Responsible-AI, model governance, and auditcompliance frameworks.

Prior enterprise-scale AI platform operations experience.

Experience & Certifications

9 12+ years in MLAI platform operations, SRE, or LLMOps with production agenticGenAI exposure (Tier 5 6).

Google Cloud Professional (MLDevOps) certification preferred.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90911958
  • Position Id: 9080267
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Santa Clara, California

23d ago

Easy Apply

Full-time

Depends on Experience

Milpitas, California

Today

Easy Apply

Contract

Palo Alto, California

Today

Full-time

USD 400,000.00 - 550,000.00 per year

Palo Alto, California

Today

Full-time

USD 350,000.00 - 500,000.00 per year

Search all similar jobs