Position: Incident Manager / Operations Specialist
Location: Plano, TX/Richmond, VA/McLean, VA/Chicago, IL
Duration: Long term contract
Tools: ServiceNow / Jira / Splunk / Dynatrace / Datadog / Grafana
Cloud: AWS / Azure / Google Cloud Platform
**This position is open to Only Ex-CapitalOne Employees/Contractors**
Job Summary:
We are seeking experienced Incident Managers / Operations Specialists responsible for managing production incidents, coordinating technical teams, monitoring applications and infrastructure, and driving timely resolution of critical issues. The ideal candidate should have strong experience in ITIL-based incident management, monitoring, escalation, root cause analysis, and operational support.
Key Responsibilities
- Manage P1/P2/P3 incidents from identification through resolution and closure.
- Coordinate technical teams, application owners, infrastructure teams, vendors, and business stakeholders during incidents.
- Monitor applications, systems, infrastructure, and services to proactively identify issues.
- Lead incident bridges and ensure effective communication throughout the incident lifecycle.
- Track incidents against SLA, response, and resolution targets.
- Perform initial troubleshooting, impact assessment, escalation, and resolution coordination.
- Conduct Root Cause Analysis (RCA) and document corrective and preventive actions.
- Identify recurring incidents and work with engineering teams to eliminate underlying issues.
- Maintain incident reports, dashboards, operational metrics, and management communications.
- Support problem management, change management, and continuous improvement initiatives.
- Participate in on-call and after-hours support as required.
Required Skills
- 10+ years of experience in Incident Management, IT Operations, Production Support, or related roles.
- Strong knowledge of ITIL Incident, Problem, and Change Management processes.
- Experience managing critical production incidents and major incident response.
- Strong knowledge of application and infrastructure monitoring and alerting.
- Experience with ticketing tools such as ServiceNow, Jira, or similar platforms.
- Strong understanding of SLA management, escalation procedures, and incident prioritization.
- Excellent troubleshooting, communication, coordination, and stakeholder management skills.
- Experience preparing RCA reports, incident metrics, and management dashboards.
Preferred Skills
- Experience with monitoring tools such as Splunk, Dynatrace, AppDynamics, Datadog, New Relic, or Grafana.
- Knowledge of AWS, Azure, or Google Cloud Platform environments.
- Understanding of cloud infrastructure, databases, APIs, networking, and application architecture.
- Experience with DevOps, CI/CD, automation, and observability.
- ITIL certification is preferred.
- Experience working in 24x7 production support / on-call environments.