Major Incident Manager

Hybrid in Phoenix, AZ, US • Posted 1 day ago • Updated 1 day ago
Full Time
75% Travel Required
Hybrid
100+
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Splunk
  • Dynatrace

Summary

We are looking for an experienced Major Incident Manager to lead the response to critical production incidents across our enterprise technology landscape. You will own the incident lifecycle end to end: command and coordination during the bridge, executive communication, root cause analysis, and the improvements that prevent repeat outages. The ideal candidate stays calm under pressure, makes sound decisions quickly, and can align technical teams and senior stakeholders around one goal: fast, safe service restoration.

Key Responsibilities

  • Lead and coordinate major incident bridges, acting as incident commander from detection to resolution.
  • Assess incident severity and business impact, and set priority in line with ITIL-aligned processes.
  • Drive escalation and decision-making during high-pressure situations, engaging the right technical and vendor teams.
  • Give timely, clear updates to executives and business stakeholders throughout the incident.
  • Use observability tools (Splunk, Dynatrace, AppDynamics, etc.) to support triage, diagnosis, and validation of recovery.
  • Lead post-incident reviews and root cause analysis (RCA), and track corrective and preventive actions to closure.
  • Work with Problem, Change, and Release Management to reduce repeat incidents and change-related failures.
  • Apply SRE principles (SLAs/SLOs, error budgets, toil reduction, automation) to improve reliability and operational resilience.
  • Maintain incident playbooks, runbooks, and technical documentation, and produce regular reporting on incident trends and KPIs (MTTR, MTTD, recurrence).
  • Support audit, compliance, and business continuity/disaster recovery requirements.

Required Skills

  • Major incident management and incident command leadership
  • Enterprise production support and operations management
  • Critical incident response and service restoration
  • Observability platforms (Splunk, Dynatrace, AppDynamics, etc.)
  • Executive and stakeholder communication
  • Incident severity assessment and business impact analysis
  • ITIL incident management processes and governance
  • Root cause analysis and post-incident reviews
  • Cross-functional technical team coordination
  • SRE concepts, SLA/SLO awareness, automation and toil reduction

Preferred Skills

  • Problem Management and Change Management
  • Monitoring and alerting tools
  • Cloud infrastructure (AWS, Azure, Google Cloud Platform)
  • Application support and middleware technologies
  • Audit and compliance management
  • Business continuity and disaster recovery
  • Process improvement and operational excellence
  • Technical documentation and reporting
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91141690
  • Position Id: 9097980
  • Posted 1 day ago
Contact the job poster
JK

Jay Kumar

Recruiter @ SN Cloud Solutions
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Phoenix, Arizona

•

Today

Easy Apply

Full-time

Depends on Experience

Chandler, Arizona

•

2d ago

Easy Apply

Full-time

80,000 - 90,000

Remote

•

Today

Full-time

USD 74,000.00 - 87,000.00 per year

North Carolina

•

Today

Full-time

Search all similar jobs