People at Apple don't just build products, they craft the kind of experiences that have revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here!
Join Apple, and help us leave the world better than we found it.
The Apple Services Engineering (ASE) team builds and provides systems and infrastructure that fuel Apple's services (such as iCloud, iTunes, Siri, and Maps). We are the foundation on which Apple's software developers build the products that our customers love.
Apple's observability and monitoring platforms are the nervous system behind the reliability of Apple's cloud services, giving thousands of engineers the visibility they need to detect, diagnose, and resolve issues before they impact customers. Our cloud monitoring platform analyzes billions of metrics per minute and is the first place incident responders turn when something goes wrong, regardless of system scale or complexity.
We're looking for a senior SRE leader to own and evolve this platform: the metrics, logging, tracing, and alerting infrastructure that underpins operational excellence across Apple Services Engineering.
Description
You'll set technical direction for reliability and operational excellence while mentoring engineers, driving automation, and partnering closely with software, infrastructure, finance, and product teams to uphold the reliability of the platform whilst shipping improvements that matter at Apple scale.
This is a senior leadership position that defines where our observability platform reliability, scalability and performance goes next, how our SRE practice evolves, including how AI reshapes it, and how we build the team and partnerships to get there. You will lead engineers solving reliability and scale problems few organizations encounter, integrating monitoring seamlessly across disparate infrastructures, hardware, software, application, and network layers, at a scale built to reach every user on the planet. You'll build a team culture that makes SRE sustainable, rewarding, and central to how Apple ships services.
The successful candidate has a strong aptitude for both technical leadership and people management, with the ability to context-switch between strategic planning and tactical execution. You should be comfortable building and scaling teams, driving complex cross-functional initiatives, and thriving under pressure, while creating an inclusive, high-trust team culture where engineers do their best work.
We believe AI will fundamentally reshape how SRE is practiced, from anomaly detection and root-cause analysis to capacity planning and toil elimination, and we're looking for a leader who shares that conviction and can drive that transformation across the organization.
Minimum Qualifications
5+ years of engineering management experience leading SRE, infrastructure, or observability/monitoring teams
Experience hiring and leading engineers, with a desire to build, grow, and mentor a team
Deep understanding of observability systems and practices: metrics, logging, tracing, alerting, SLOs, error budgets, and fault analysis at scale
Strong systems background, comfortable troubleshooting across the full stack (network, OS, container runtime, application)
Experience operating large-scale, multi-tenant distributed systems in production, including Kubernetes environments
Practical, solid knowledge of shell/bash scripting and at least one higher-level production language (Python preferred; Go, Java, or Scala also valued)
Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operations
Track record of building high-performing teams through coaching, clear expectations, and psychological safety
Demonstrated ability to drive cross-functional initiatives to completion and communicate at the executive level
Bachelor's or Master's degree in Computer Science, Engineering, or related field, or equivalent experience
Preferred Qualifications
Deep familiarity with the Prometheus ecosystem and cloud-native observability stacks (Thanos, Splunk, OpenTelemetry, or similar)
Experience with third-party cloud platforms (AWS, Google Cloud Platform, or Azure) and infrastructure as code (Terraform, Ansible)
Comfortable with open-source configuration management and orchestration tools (Helm, Puppet, Spinnaker)
Demonstrable knowledge of TCP/IP, HTTP, web application security, and multi-tier web application architectures
Experience running infrastructure as an internal managed service with defined SLAs
Familiarity with microservices architecture and container orchestration with Kubernetes at scale
Background in capacity planning, performance engineering, or infrastructure architecture
Track record of driving cultural and process transformation within SRE organizations
Experience building or deploying AI-powered operational tooling (AIOps, intelligent alerting, automated diagnostics)
Developing and delivering multi-mode communications tailored to the unique needs of different audiences
Anticipating and balancing the needs of multiple stakeholders
Making sense of complex, high-quantity, and sometimes contradictory information to solve problems effectively
Rebounding from setbacks and adversity when facing difficult situations
Knowing the most effective and efficient processes to get things done, with a focus on continuous improvement
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 90733111
- Position Id: bf98c540d07389652e32380c67bdeda8
- Posted 1 hour ago