NOC / Service Management Analyst
Location: Richardson, TX
Work Arrangement: On-Site
Job Summary
We are seeking an experienced NOC / Service Management Analyst to join our operations team in Richardson, TX. This is an on-site position responsible for monitoring enterprise network and system health, responding to alerts, managing incidents, and coordinating technical resources during service disruptions.
The ideal candidate will have experience working in a Network Operations Center (NOC), IT Service Management (ITSM), or enterprise operations environment, with strong knowledge of ServiceNow, incident management, monitoring tools, and major incident coordination.
This role serves as a critical link between monitoring operations, technical support teams, vendors, and business stakeholders to help maintain service availability and operational continuity.
Key Responsibilities
Monitoring & Alert Response
- Monitor enterprise network, infrastructure, applications, and system health using monitoring and event-management tools.
- Respond to system alerts, alarms, and operational events in accordance with established procedures.
- Perform initial triage and troubleshooting of incidents.
- Collaborate with Level 2 and Level 3 technical support teams to investigate and resolve issues.
- Escalate incidents based on severity, business impact, service-level requirements, and established escalation procedures.
- Identify recurring alerts and operational issues and recommend improvements to monitoring and response procedures.
Incident & Major Incident Management
- Create, update, and maintain accurate incident records in ServiceNow, serving as the system of record.
- Ensure incident tickets contain complete and timely information, including impact, troubleshooting actions, escalation details, and resolution status.
- Serve as Incident Manager during technical bridge calls for major incidents and service outages.
- Coordinate L2/L3 support teams, application teams, infrastructure teams, vendors, and other technical resources during major incidents.
- Establish and maintain clear communication throughout the incident lifecycle.
- Provide timely status updates to appropriate technical and business stakeholders.
- Track incident progress through resolution and ensure appropriate closure documentation is completed.
- Support post-incident reviews and identify opportunities to improve incident response and service reliability.
Knowledge Management
- Create, maintain, and update Knowledge Base (KB) articles for recurring alerts, incidents, troubleshooting procedures, and operational processes.
- Ensure knowledge articles contain accurate and actionable handling instructions.
- Review existing KB documentation periodically and update procedures as systems, applications, or operational processes change.
- Help ensure operational knowledge is accessible to NOC and technical support teams.
Application & Batch Operations
- Verify startup and shutdown procedures for critical applications and services.
- Monitor and confirm successful completion of critical application jobs, batch processes, and operational KPIs.
- Investigate failed or delayed jobs and coordinate escalation with appropriate technical teams.
- Follow documented operational runbooks and procedures for scheduled and event-driven activities.
Vendor & Change Coordination
- Communicate vendor-related outages, service interruptions, and technical issues to appropriate L2/L3 support teams.
- Coordinate with vendors and internal technical teams during vendor-related incidents.
- Support change-related activities and communicate potential service impacts to appropriate stakeholders.
- Ensure operational procedures and incident documentation reflect approved changes where applicable.
Required Qualifications
- Experience in a NOC, Network Operations, IT Service Management, IT Operations, or Enterprise Operations environment.
- Hands-on experience with ServiceNow Incident Management or a comparable ITSM platform.
- Experience monitoring enterprise network, infrastructure, applications, or systems.
- Understanding of Incident Management, Major Incident Management, escalation procedures, and ITIL/ITSM concepts.
- Experience coordinating technical teams during outages or major incidents.
- Strong troubleshooting, analytical, and problem-solving skills.
- Excellent written and verbal communication skills.
- Ability to work effectively in a high-pressure environment and manage multiple incidents simultaneously.
- Ability to follow operational procedures, runbooks, and escalation matrices.
- Willingness and ability to work on-site in Richardson, TX.
Preferred Qualifications
- ITIL Foundation certification or equivalent ITSM knowledge.
- Experience with enterprise monitoring tools such as SolarWinds, Splunk, Dynatrace, AppDynamics, Nagios, or similar platforms.
- Experience supporting network infrastructure, servers, applications, or enterprise systems.
- Experience with major incident bridge management and stakeholder communications.
- Knowledge of batch scheduling and enterprise job-monitoring processes.
- Experience working with external technology vendors and service providers.
- Experience in a 24x7 NOC or operations environment.