Senior AIOps and Incident Management / Site Reliability Engineering

Hybrid in Fort Mill, SC, US • Posted 4 days ago • Updated 4 days ago
Contract W2
12 Months
No Travel Required
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

📋 Comparing job requirements...

Job Details

Skills

  • DYNATRACE
  • AI
  • SRE
  • ITSM
  • ITIL
  • DEVOPS
  • AWS

Summary

Senior AIOps and Incident Management / Site Reliability Engineering

We are seeking a Senior AIOps Incident Manager & Site Reliability Engineer to lead incident management, operational resilience, and intelligent automation initiatives across enterprise technology environments. This role will partner with Network Operations Center (NOC), Infrastructure Operations, Cloud Engineering, DevOps, and Application Support teams to proactively detect, respond to, and prevent technology incidents.

The ideal candidate combines hands-on incident management expertise with experience implementing observability, automation, and AI-driven operational solutions to improve system reliability, reduce operational overhead, and enhance customer experience. The candidate should possess deep expertise in AIOps, ITSM, ITIL, SRE, Incident Management, Cloud Operations, and Enterprise Infrastructure.

This is a hybrid role requiring on-site attendance up to three days per week. Candidates must reside within approximately one hour commuting distance of one of the following office locations:

  • Fort Mill, SC
  • Austin, TX
  • Boston, MA
  • New York, NY
  • Tempe, AZ
  • San Diego, CA

Incident & Recovery Management

  • Monitor, document, and analyze major incident response efforts and service recovery activities.
  • Serve as a senior escalation point for Tier 1 and Tier 2 operational incidents.
  • Conduct incident reviews, root cause analysis, and corrective action planning.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).

Site Reliability Engineering

  • Implement SRE practices to improve platform reliability, scalability, and resiliency.
  • Define and monitor SLAs, SLOs, and operational KPIs.
  • Develop proactive reliability and availability strategies.

AIOps & Automation

  • Implement AIOps solutions to automate incident detection, diagnosis, remediation, and prevention.
  • Build and optimize AI-powered operational agents and self-healing workflows.
  • Reduce operational effort through intelligent automation.

Observability & Monitoring

  • Lead enterprise monitoring initiatives using Dynatrace and related observability platforms.
  • Improve visibility across cloud, infrastructure, applications, and user experiences.
  • Enable predictive monitoring and anomaly detection.

ITSM & Service Operations

  • Develop and enhance incident, problem, change, and event management frameworks aligned with ITIL and ITSM best practices.
  • Leverage ServiceNow workflow automation to improve service delivery.

Cross-Functional Leadership

  • Partner with Infrastructure, DevOps, Cloud, Security, Application Development, and NOC teams.
  • Mentor operational teams and promote an automation-first culture.

Qualifications

  • Bachelor''s degree in Computer Science, Information Technology, Engineering, or related field (or equivalent experience).
  • 8+ years of experience in IT Operations, Site Reliability Engineering, Infrastructure Operations, Network Operations, or Production Support environments.
  • 5+ years of experience leading incident management, operational transformation, or reliability engineering initiatives.

Strong experience with:

  • Site Reliability Engineering (SRE)
  • IT Service Management (ITSM)
  • ITIL Framework
  • Incident, Problem, Change, and Event Management
  • Network Operations Center (NOC)
  • Infrastructure Operations
  • Service Desk Operations
  • Application Production Support
  • Cloud Platforms (AWS, Azure, or Google Cloud Platform)
  • DevOps practices and toolchains

Preferred Experience

  • Hands-on experience with Dynatrace, monitoring platforms, and observability solutions.
  • Experience using ServiceNow for ticketing, workflow automation, and service management.
  • Strong understanding of infrastructure, networking, cloud architecture, and enterprise application ecosystems.
  • Proven experience conducting root cause analysis and implementing preventive controls.
  • Experience leading enterprise AIOps implementations.
  • Experience building AI-powered operational agents and intelligent automation solutions.

Preferred Certifications

  • ITIL Foundation or ITIL Managing Professional
  • Certified Site Reliability Engineer (SRE)
  • AWS, Azure, or Google Cloud certifications
  • ServiceNow certifications

Nice to Have

  • Experience with workflow orchestration and enterprise automation platforms.
  • Familiarity with predictive analytics, machine learning operations, and autonomous operations frameworks.

 

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10120317
  • Position Id: 9029227
  • Posted 4 days ago
Contact the job poster
Chandee Pandey

Chandee Pandey

Recruiter! @ Sovereign Technologies
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Fort Mill, South Carolina

Yesterday

Easy Apply

Contract, Third Party

Depends on Experience

Charlotte, North Carolina

Today

Full-time

USD 160,000.00 - 180,000.00 per year

Fort Mill, South Carolina

Today

Easy Apply

Full-time

$52-58/hr.

Charlotte, North Carolina

27d ago

Easy Apply

Contract, Third Party

Depends on Experience

Search all similar jobs