Key Responsibilities
Production Support & Incident Management
- Lead production support operations across multiple applications and services.
- Establish and manage incident management processes, ensuring timely triage, escalation, resolution, and communication.
- Drive Major Incident Management (MIM) activities and coordinate war room sessions during critical outages.
- Conduct root cause analysis (RCA) reviews and track corrective and preventive actions.
- Monitor incident trends and identify improvement opportunities to reduce recurring issues.
- Ensure adherence to operational support procedures and best practices.
Service Level Agreement (SLA) Management
- Define, monitor, and govern operational SLAs, OLAs, and KPIs.
- Track service performance against agreed targets and drive remediation plans for SLA breaches.
- Develop and implement service improvement plans to enhance customer satisfaction and operational efficiency.
- Provide regular SLA performance reports to leadership and clients.
Technical Debt Management
- Establish governance processes for identification and prioritization of technical debt.
- Partner with engineering and architecture teams to create remediation roadmaps.
- Track technical debt backlog and ensure alignment with business priorities.
- Drive continuous improvements to system stability, maintainability, and performance.
Security Vulnerability & Risk Management
- Coordinate vulnerability assessment, remediation, and compliance activities across application and infrastructure teams.
- Track security findings from scans, audits, and assessments.
- Ensure timely resolution of critical and high-risk vulnerabilities.
- Collaborate with cybersecurity teams to support compliance, audit readiness, and risk mitigation initiatives.
- Monitor operational risks and develop mitigation strategies.
Metrics, Reporting & Governance
- Define and maintain operational dashboards and executive-level reporting.
- Analyze performance metrics including:
- Incident volumes
- Mean Time to Acknowledge (MTTA)
- Mean Time to Resolve (MTTR)
- SLA adherence
- Availability and uptime
- Technical debt reduction
- Vulnerability remediation status
- Release quality metrics
- Present operational insights and recommendations to leadership.
- Establish governance routines, review meetings, and performance tracking mechanisms.
Client & Stakeholder Management
- Act as the primary point of contact for operational and production support matters.
- Build strong relationships with business stakeholders and clients.
- Communicate service health, risks, escalations, and improvement initiatives.
- Facilitate regular operational reviews and executive status meetings.
- Manage customer expectations and drive issue resolution to completion.
Cross-Team & Vendor Coordination
- Coordinate activities across development, infrastructure, security, QA, DevOps, and support teams.
- Manage relationships with third-party vendors and managed service providers.
- Ensure vendor accountability for service commitments and deliverables.
- Lead cross-functional dependencies and facilitate issue resolution across organizational boundaries.
- Drive alignment among geographically distributed and matrixed teams.
Continuous Improvement & Operational Excellence
- Identify process improvement opportunities and implement best practices.
- Support automation initiatives that improve service reliability and operational efficiency.
- Drive ITIL-aligned service management practices.
- Lead post-incident reviews and ensure lessons learned are incorporated into future processes.
- Foster a culture of accountability, continuous improvement, and customer focus.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, Business, or related field
- 8+ years of experience in IT Program Management, Service Delivery, Production Support, or Operations Management.Strong experience managing large-scale production support environments.Proven experience with incident management, problem management, and change management processes
- 8+ years of Experience managing multiple cross-functional teams and vendors.Strong understanding of application support, cloud platforms, infrastructure, and software delivery lifecycles.
- 8+ years of Experience with operational dashboards, KPI reporting, and executive communications.Excellent stakeholder management and client-facing communication skills.
Preferred Qualifications
- ITIL Foundation or ITIL Managing Professional certification.
- Experience with Agile, Scrum, DevOps, and Site Reliability Engineering (SRE) practices.
- Familiarity with ServiceNow, AppDynamics, or other monitoring tools.
- Experience in regulated environments and security compliance frameworks.
Akanksha Agarwal | Sr. US IT Recruiter