Title: Senior QA Automation Engineer
Location: Phoenix, AZ (Hybrid)
Work Stream - End to end Observability
Skills:
Onsite Senior QA Automation Engineer / QA Lead with Banking Payments, API, Observability, Reliability, and Production Support testing experience.
Job Summary:
We are seeking an experienced Engineering Manager – Observability & Monitoring to lead engineering teams responsible for enterprise monitoring, observability, reliability, and production operations. The ideal candidate will have strong experience across metrics, logs, traces, APM, cloud monitoring, SRE, and automation, with proven technical and people leadership skills.
Key Responsibilities:
- Lead and mentor engineering teams responsible for Observability, Monitoring, APM, Logging, and Alerting platforms.
- Define and implement enterprise-wide observability strategy, standards, and engineering practices.
- Build and manage solutions for metrics, logs, distributed tracing, application performance, infrastructure monitoring, and real-time alerting.
- Drive adoption of SRE practices, SLIs, SLOs, SLAs, error budgets, and reliability engineering.
- Lead implementation and optimization of platforms such as Datadog, Splunk, Grafana, Prometheus, OpenTelemetry, CloudWatch, and Azure Monitor.
- Develop automated monitoring, anomaly detection, event correlation, alert remediation, and self-healing capabilities.
- Partner with Cloud, Platform Engineering, DevOps, Application Engineering, and Security teams.
- Establish dashboards, operational KPIs, alerting standards, and production health metrics.
- Lead incident management, root-cause analysis, problem management, and continuous reliability improvements.
- Drive observability automation through Python/Java/Node.js, APIs, CI/CD, Terraform, and infrastructure automation.
- Manage engineering roadmaps, technical delivery, budgets, vendor relationships, and team development.
Required Skills:
- 10+ years of experience in software engineering, SRE, DevOps, Platform Engineering, or Observability.
- 3+ years of engineering management or technical leadership experience.
- Strong hands-on experience with Monitoring & Observability platforms.
- Expertise in Datadog, Splunk, Prometheus, Grafana, OpenTelemetry, and/or CloudWatch.
- Strong knowledge of AWS/Azure/Google Cloud Platform, Kubernetes, Docker, Linux, and cloud-native architectures.
- Experience with APM, distributed tracing, centralized logging, metrics, alerting, and event management.
- Strong understanding of SRE, reliability engineering, incident management, and production operations.
- Experience with Python, Java, Node.js, REST APIs, automation, and CI/CD.
- Excellent people leadership, stakeholder management, communication, and problem-solving skills.