AI OPS Engineer

Remote in Denver, CO, US • Posted 1 hour ago • Updated 1 hour ago
Contract W2
On-site
Fitment

Dice Job Match Score™

🤯 Applying directly to the forehead...

Job Details

Skills

  • Generative Artificial Intelligence (AI)
  • Software Engineering
  • Python
  • API
  • Data Engineering
  • Bridging
  • IT Operations
  • FOCUS
  • Dashboard
  • Reasoning
  • Data Integration
  • Root Cause Analysis
  • Orchestration
  • Microsoft Certified Professional
  • Real-time
  • Workflow
  • Configuration Management Database
  • Google Cloud
  • Google Cloud Platform
  • ServiceNow
  • Zabbix
  • Dynatrace
  • Enterprise Search
  • Management
  • Return On Investment
  • Use Cases
  • Help Desk
  • Incident Management
  • Marketing Intelligence
  • MI
  • Microsoft Windows
  • Collaboration
  • MEAN Stack
  • Large Language Models (LLMs)
  • Prompt Engineering
  • Artificial Intelligence

Summary

AI OPS Engineer II

While modern generative AI and LLM orchestration tools have only been mainstream for a short time, the 3 to 5 years of required experience for this Level II role reflects the total engineering maturity needed for enterprise deployment. Candidates are expected to have 1 to 2 years of direct, hands-on experience with modern LLMs, multi-agent frameworks, and AI coding assistants, built on top of a solid 2 to 3-year foundational background in backend software engineering (Python), API integration, data engineering, or traditional AIOps. This combined 3-to-5-year tenure ensures the candidate not only understands how to orchestrate AI workflows, but also possesses the established systems thinking required to securely and reliably integrate these tools into complex, enterprise-grade infrastructure platforms like ServiceNow, Dynatrace, and Google Cloud Platform.The AI Automation Engineer will be the foundational builder for our Enterprise Operational AI Platform, bridging the gap between advanced multi-agent AI architecture and our core IT operations. Their initial focus will be designing robust data integrations that connect our LLM framework to essential infrastructure systems like ServiceNow, Dynatrace, Zabbix, and Google Cloud Platform. Using tools like the Model Context Protocol (MCP) and AI-assisted development (Devin/Windsurf), they will execute a phased rollout strategy starting with read-only workflows for incident triage and root cause analysis (RCA), and progressively advancing toward fully autonomous Help Desk automation and proactive self-healing capabilities. Ultimately, their success will be measured by improved AI observability, reduced mean time to resolution (MTTR), and the seamless orchestration of domain-specific AI agents that provide a unified, intelligent interface for our I&O engineers.

Our Operations team is seeking an innovative and highly technical AI Automation Engineer to help build and scale our unified operational AI interface. This role is completely focused on the design, data integration, and agentic framework of our enterprise AI platform, which serves as the single, centralized entry point for I&O engineers to access operational data and insights.
You will partner closely with our lead AI architects to operationalize this vision. Your work will center on building the foundational data integrations, developing multi-agent reasoning workflows, and implementing AI-driven use cases for Help Desk automation, incident triage, and root cause analysis (RCA).

About the Operational AI Platform & Architecture
Our Enterprise Operational AI Platform is DaVita s unified operational AI interface. It acts as the centralized entry point for I&O engineers to access enterprise-wide operational data, insights, and automation capabilities eliminating the need to navigate fragmented tools and dashboards.
From an architectural standpoint, the platform is designed as a multi-agent reasoning system powered by advanced LLMs. Key architectural components include:

Skills:

Agentic Framework & MCP: It utilizes a dynamic agent registry and the Model Context Protocol (MCP) to orchestrate real-time, interactive operational tasks across different domain-specific AI agents.
Foundational Data Integration: The platform integrates directly with I&O systems of record and telemetry (e.g., ServiceNow CMDB, Dynatrace, Zabbix, Google Cloud Platform) to ingest infrastructure state, logs, and monitoring data.
Phased Execution Model: The architecture supports a scalable adoption path, beginning with query-only/read-only contextualization (e.g., incident triage, RCA) and advancing into fully autonomous, agentic actions (e.g., proactive outage prevention, maintenance suppression, and coordinated self-healing).
Complementary Enterprise Search: It is architected to work alongside and cross-validate data with enterprise search tools (like Glean) while providing the dynamic orchestration capabilities required for real-time operations.
Key Responsibilities
Operational AI Interface Development
Assist in the architecture and development of the unified AI assistant, ensuring a seamless experience for I&O engineers.
Develop and register AI agents and Model Context Protocol (MCP) integrations to handle real-time, interactive operational tasks.
Build backend workflows that support discovery, platform engineering sizing decisions, and proactive outage prevention.
Foundational Data & Integration
Develop robust data integrations connecting the AI platform to foundational infrastructure sources, including CMDB, Google Cloud Platform, ServiceNow, Zabbix, and Dynatrace.
Automate the ingestion and contextualization of infrastructure data across the entire I&O organization, leveraging AI-assisted development tools like Devin and Windsurf to accelerate engineering and accurately perform reactive and scheduled operational tasks.
Collaborate with observability and enterprise search (e.g., Glean) teams to cross-validate data, manage knowledge quality, and prevent LLM hallucinations.
Use Case Delivery & ROI
Deliver initial "low-hanging fruit" use cases, specifically focusing on Help Desk automation and Incident Management (MI) support.
Design AI capabilities to assist with change windows, maintenance suppression, and coordination of self-healing actions.
Establish AI observability metrics to measure confidence levels, token usage, and improvements to Mean Time to Resolution (MTTR).

Required Skill Sets & Qualifications
Technical Proficiencies:
AI & LLM Development: Deep understanding of Large Language Models (LLMs), advanced prompt engineering, and Retrieval-Augmented Generation (RAG).
AI Developer Tools: Hands-on experience or familiarity with advanced AI coding assistants and autonomous engineering platforms such as Devin and Windsurf.
Agentic Frameworks: Hands-on expe
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: cxbcsi
  • Position Id: Job44923
  • Posted 1 hour ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Denver, Colorado

Today

Easy Apply

Contract

Depends on Experience

Denver, Colorado

Today

Easy Apply

Contract

Depends on Experience

Hybrid in Littleton, Colorado

5d ago

Easy Apply

Contract

Depends on Experience

Remote

Today

Easy Apply

Contract

Depends on Experience

Search all similar jobs