Engineering Manager – AI Ops & Autonomous Operations Platform We are seeking an Engineering Manager to build and lead an enterprise AI Ops platform that transforms IT operations through AI, automation, and observability.
Lead the development of a greenfield AI Ops and Autonomous Operations platform. Build capabilities for proactive monitoring, intelligent alerting, and self-healing systems. Establish enterprise standards for observability, reliability, and operational intelligence. Own the AIOps roadmap, architecture, and delivery strategy.
Lead teams of SREs, AI Engineers, Platform Engineers, and Automation Engineers. Implement AI-driven event correlation and root cause analysis. Drive automation of incident triage, diagnosis, and remediation workflows. Develop GenAI and Agentic AI solutions for operational support.
Leverage telemetry data from logs, metrics, traces, events, and CMDB. Build enterprise observability solutions using OpenTelemetry and modern monitoring platforms. Define and mature SLI, SLO, SLA, and error-budget frameworks. Improve platform reliability, resilience, and operational efficiency.
Reduce alert noise and operational toil through intelligent automation. Partner with Infrastructure, Cloud, Security, Risk, and Application teams. Lead major incident management and post-incident improvement programs. Establish AI-assisted operations practices and governance models. Drive cost optimization, capacity forecasting, and predictive operations. Deliver executive dashboards aligned to business outcomes and customer experience. Build a self-service operational intelligence platform for engineering teams. Create measurable improvements in MTTR, MTTA, availability, and engineering productivity.
Key Skills: AIOps, SRE, Observability, OpenTelemetry, Datadog, Splunk, ServiceNow, AI/ML, GenAI, Agentic AI, Kubernetes, Cloud Platforms, Incident Management, Automation, Platform Engineering, Reliability Engineering.