Engineering Manager - AIOps

Hybrid in Phoenix, AZ, US • Posted 1 day ago • Updated 1 day ago
Contract W2
12 Months
No Travel Required
Hybrid
$65 - $70/hr
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Generative Artificial Intelligence (AI)
  • IT Operations
  • Incident Management
  • Kubernetes
  • Machine Learning (ML)
  • Operational Efficiency
  • Optimization
  • Productivity
  • Reliability Engineering
  • Roadmaps
  • Root Cause Analysis
  • ServiceNow
  • Workflow
  • Customer Experience
  • Dashboard
  • Forecasting
  • Splunk
  • Artificial Intelligence
  • Budget
  • Cloud Computing
  • Cloud Security
  • Configuration Management Database
  • SLA
  • observability
  • SRE

Summary

Role: Engineering Manager 

Location: Phoenix, AZ (Hybrid)

Work Stream - AI Ops

Engineering Manager – AI Ops & Autonomous Operations Platform 

We are seeking an Engineering Manager to build and lead an enterprise AI Ops platform that transforms IT operations through AI, automation, and observability.

· Lead the development of a greenfield AI Ops and Autonomous Operations platform.

· Build capabilities for proactive monitoring, intelligent alerting, and self-healing systems.

· Establish enterprise standards for observability, reliability, and operational intelligence.

· Own the AIOps roadmap, architecture, and delivery strategy.

· Lead teams of SREs, AI Engineers, Platform Engineers, and Automation Engineers.

· Implement AI-driven event correlation and root cause analysis.

· Drive automation of incident triage, diagnosis, and remediation workflows.

· Develop GenAI and Agentic AI solutions for operational support.

· Leverage telemetry data from logs, metrics, traces, events, and CMDB.

· Build enterprise observability solutions using OpenTelemetry and modern monitoring platforms.

· Define and mature SLI, SLO, SLA, and error-budget frameworks.

· Improve platform reliability, resilience, and operational efficiency.

· Reduce alert noise and operational toil through intelligent automation.

· Partner with Infrastructure, Cloud, Security, Risk, and Application teams.

· Lead major incident management and post-incident improvement programs.

· Establish AI-assisted operations practices and governance models.

· Drive cost optimization, capacity forecasting, and predictive operations.

· Deliver executive dashboards aligned to business outcomes and customer experience.

· Build a self-service operational intelligence platform for engineering teams.

· Create measurable improvements in MTTR, MTTA, availability, and engineering productivity.

Key Skills: AIOps, SRE, Observability, OpenTelemetry, Datadog, Splunk, ServiceNow, AI/ML, GenAI, Agentic AI, Kubernetes, Cloud Platforms, Incident Management, Automation, Platform Engineering, Reliability Engineering.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91170837
  • Position Id: 9070273
  • Posted 1 day ago

Company Info

About TechVirtue LLC

TechVirtue is involved in developing a wide range of solutions in finding the perfect candidate who has a strong knowledge in his/her work and suits the company's work culture. We even provide one-stop solutions ranging from software development and maintenance to expert support and advisory. Our team consists of experts who have several years of experience in staffing, recruitment, and web development. Our dedicated and motivated team makes sure to fulfill all our customers requirements.

About_Company_OneAbout_Company_Two
Contact the job poster
NK

Nikhil Kanchi

Recruiter @ TechVirtue LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs