Location: Cincinnati, OH (100% Onsite)
Travel: Occasional travel required
On-Call: Participation in a rotating on-call schedule required
Job Overview:
o We are seeking a Major Incident & Site Reliability Engineer to support critical production environments by leading high-severity incident response, driving service restoration, and improving system reliability.
o This role requires strong technical troubleshooting skills, cross-functional collaboration, and experience managing enterprise production incidents in a 24x7 environment.
Responsibilities:
o Lead P1/P2 major incident response, coordinating cross-functional teams to restore critical services.
o Facilitate incident bridge calls, provide executive communications, and drive timely resolution.
o Perform Root Cause Analysis (RCA) and partner with engineering teams to implement corrective actions.
o Troubleshoot production issues across applications, infrastructure, databases, and cloud environments.
o Support production operations, monitoring, deployments, and operational readiness.
o Participate in on-call rotations and collaborate with development, infrastructure, and operations teams to improve platform reliability.
Required Skills:
o Experience leading Major Incident Management (P1/P2) in a 24x7 enterprise environment.
o Strong background in Site Reliability Engineering (SRE) or L3 Production Support.
o Experience with ServiceNow, ITIL, and enterprise monitoring tools such as Splunk, Grafana, Dynatrace, or similar.
o Hands-on troubleshooting experience with Linux/Unix, cloud platforms, databases, and distributed applications.
o Scripting experience with Python, Bash, or PowerShell is preferred.
o Excellent communication skills with the ability to coordinate technical and business stakeholders during critical incidents.