Senior L2 Production Support Engineer – Video Streaming / OTT
Position Summary
We are seeking a Senior L2 Production Support Engineer to support the stability, availability, and operational performance of 24/7 live video streaming, advertising, player, and real-time delivery platforms.
The engineer will take end-to-end ownership of customer-impacting production incidents escalated from L1 Support. This is a hands-on, customer-facing role involving production troubleshooting, incident management, automation, Infrastructure as Code, and operational improvements. The role will also serve as a key operational bridge between Support, Engineering, DevOps, partners, and customers, particularly during high-impact live events.
Key Responsibilities
Incident & Production Support
- Own and drive escalated customer issues from L1 Support through resolution.
- Troubleshoot complex production incidents affecting live streaming, VOD playback, ad insertion, DRM, and WebRTC services.
- Work directly within production environments to perform approved configuration changes, CDN adjustments, mitigations, and emergency changes when required.
- Lead or actively participate in live incident bridges involving customers, internal teams, and external partners.
- Provide clear and timely incident updates, including customer-facing communications.
- Monitor production systems and respond to critical alerts in accordance with established SLAs.
Infrastructure as Code & Operations
- Work with Infrastructure as Code (IaC) to troubleshoot and safely modify production environments.
- Utilize technologies and workflows such as Terraform, Helm, Kubernetes manifests, GitOps, CI/CD, and deployment pipelines.
- Execute infrastructure and configuration changes through safe, auditable, and repeatable processes.
- Collaborate with Engineering and DevOps teams to improve deployment reliability and operational safety.
AI-Driven Operations & Automation
- Leverage AI tools and automation to improve operational efficiency and incident response.
- Support AI-assisted incident triage, classification, alert correlation, and automated runbook execution.
- Use AI to improve incident communications and accelerate troubleshooting.
- Identify recurring incident patterns and opportunities for automation and operational scalability.
- Contribute to automation-first and AI-augmented operational practices.
Pre-Event Readiness
- Participate in operational readiness activities for critical customer and live-streaming events.
- Validate runbooks, monitoring coverage, system readiness, and potential operational risks.
- Develop and rehearse incident response strategies for high-risk scenarios.
- Coordinate with customers and internal teams to ensure smooth event execution.
24/7 On-Call Operations
- Participate in a global 24/7 on-call rotation, including nights, weekends, and holidays.
- Respond to critical alerts related to stream health, player performance, and delivery infrastructure.
- Ensure effective handoffs between shifts and global support regions.
Root Cause Analysis & Continuous Improvement
- Perform and contribute to Root Cause Analysis (RCA) for production incidents.
- Document incident findings, corrective actions, and preventive measures.
- Identify recurring issues and partner with Engineering and Product teams to address systemic problems.
- Create and improve operational runbooks, playbooks, and knowledge-base documentation.
Engineering & Cross-Functional Collaboration
- Partner with Engineering teams to escalate defects, validate fixes, and support production deployments.
- Provide feedback on observability, monitoring, tooling gaps, and operational risks.
- Represent operational requirements during post-incident reviews.
- Collaborate with Support, Engineering, DevOps, Product, and customer-facing teams to improve overall platform reliability.
Required Qualifications & Skills
- 5+ years of experience in production operations, technical support, operational support, SRE, or a related customer-facing technical role.
- Proven ability to independently own complex technical issues from investigation through resolution.
- Strong experience supporting production video streaming, OTT, or live-streaming platforms.
- Strong troubleshooting experience across distributed systems, APIs, microservices, and cloud infrastructure.
- Knowledge of HLS, DASH, CMAF, WebRTC, DRM, and CDN architectures.
- Experience with monitoring, alerting, and log-analysis tools such as:
- Grafana
- Prometheus
- Kibana / ELK
- Loki
- Ability to correlate backend streaming metrics, player telemetry, and CDN signals to diagnose customer-impacting issues.
- Experience making controlled changes within production environments.
- Working knowledge of incident management, production operations, and on-call support.
- Strong understanding of Infrastructure as Code and deployment practices, including Terraform, Helm, Kubernetes, GitOps, and CI/CD.
Operational & Communication Skills
- Strong sense of ownership and accountability for customer outcomes.
- Ability to remain structured and decisive during high-pressure production incidents.
- Excellent written and verbal communication skills.
- Comfortable communicating technical issues and status updates directly to customers.
- Strong collaboration skills across Engineering, DevOps, Support, Product, and other cross-functional teams.
- Ability to manage multiple priorities in a fast-paced 24/7 production environment.