Role overview: We are looking for an Incident & Problem Management Coordinator who can lead production incident bridge calls, coordinate technical triage, maintain runbooks, communicate effectively with senior leadership, and drive issues through root-cause analysis and permanent resolution. Key responsibilities: - Lead major-incident bridge calls and establish clear ownership, priorities and next actions.
- Coordinate application, data, infrastructure, network, cloud, security and vendor teams during incidents.
- Assess and document business impact, affected services, customer impact, severity and restoration priorities.
- Provide accurate and timely incident communications to senior leaders and business stakeholders.
- Maintain incident timelines, decisions, actions, owners and resolution status.
- Create, maintain and execute triage runbooks for recurring and high-risk incidents.
- Track incidents from detection through containment, service restoration and formal closure.
- Coordinate root-cause analysis and prepare clear RCA reports covering cause, impact, resolution and preventive actions.
- Own and maintain the problem log, including recurring incidents, known errors, workarounds, action owners and due dates.
- Drive corrective and preventive actions to completion and escalate overdue items.
- Review monitoring and alerting coverage and identify gaps in incident detection.
- Define dashboards, alert-routing procedures, escalation matrices and operational health reports.
- Conduct post-incident reviews and identify opportunities for automation and operational improvement.
- Report incident trends, recurring problems, SLA performance and operational risks to leadership.
Mandatory skills: - Strong experience in major incident management and problem management.
- Proven ability to lead technical bridge calls during critical production incidents.
- Excellent written and verbal executive-communication skills.
- Experience creating triage runbooks, incident reports, RCA documents and problem records.
- Understanding of infrastructure, networks, applications, integrations, databases and data pipelines.
- Experience with ITSM platforms such as ServiceNow, Jira Service Management or equivalent.
- Hands-on understanding of observability, dashboards, logging, monitoring and alerting.
- Ability to remain calm, structured and decisive during high-severity incidents.
- Willingness to work in the EST time zone and support critical incidents when required.
Preferred skills: |