Site Reliability Engineer, Apple Data Platform / Big Data Platform

Austin, TX, US • Posted 1 day ago • Updated 10 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

⭐ Evaluating experience...

Job Details

Skills

  • Art
  • Music
  • System Integration Testing
  • Incident Management
  • Analytics
  • Machine Learning (ML)
  • Artificial Intelligence
  • IaaS
  • Roadmaps
  • Data Engineering
  • Computer Science
  • Reliability Engineering
  • DevOps
  • Python
  • Golang
  • Big Data
  • Apache Spark
  • Apache Flink
  • Kubernetes
  • Cloud Computing
  • Amazon Web Services
  • Google Cloud
  • Google Cloud Platform
  • Communication
  • Production Support
  • Data Governance
  • Customer Facing
  • Technical Support
  • Grafana
  • Splunk
  • Continuous Integration
  • Continuous Delivery
  • Amazon S3
  • Cloud Storage
  • Computer Networking
  • Extract
  • Transform
  • Load
  • Orchestration
  • Workflow
  • Scheduling
  • Scripting

Summary

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books - at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.

Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success - running incident response, providing hands-on support to internal teams, and partnering with developers to make cutting-edge services like Spark, Flink, Airflow, Trino, Notebooks, and LLM-based agent platforms reliable at scale.

Description

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple - while specialising in the big data engines and catalog/governance layers that power analytics and data engineering across the company. As an SRE on Apple Data Platform, you'll operate and support the team's full portfolio, from ML/AI platform services to multi-cloud infrastructure, and grow into the team's go-to expert for big data platform services - including Spark, Flink, Airflow, Trino, Notebooks, REST Catalog services (such as Glue Catalog), and data governance. Just as importantly, you'll be a first point of contact for the internal customers who rely on these services daily - someone who can translate a confusing error or a vague support request into a clear diagnosis and a fast resolution.

We're looking for a self-motivated engineer who thrives on ownership - someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love solving hard operational problems, take genuine satisfaction in helping frustrated customers get unblocked, and want a front-row seat to how Apple's data engineering platform scales, this role offers real room to grow your scope and impact over time.

Minimum Qualifications

Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.

1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.

Proficient in Python; working knowledge of Golang a plus.

Deep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).

Experience with Kubernetes and at least one major cloud provider (AWS or Google Cloud Platform).

Excellent written and verbal communication skills, with the ability to explain technical issues clearly to non-expert customers.

Solid grounding in SRE principles, with prior on-call, production-support, or customer-facing support role experience.

Preferred Qualifications

Experience with REST Catalog services (e.g., Glue Catalog) and data governance frameworks.

Prior experience in a customer-facing or technical support role, with a demonstrated passion for customer success.

Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.

Working knowledge of CI/CD pipelines and deployment workflows.

Experience with S3 and cloud storage/networking fundamentals.

Familiarity with data pipeline orchestration and workflow scheduling patterns.

A track record of automating manual operations through scripting or tooling.

Intellectual curiosity and a drive to keep learning - for yourself, your team, and the org.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90733111
  • Position Id: 63b006c758fefbb5bb66467031620b1f
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Austin, Texas

Today

Full-time

Austin, Texas

Today

Full-time

Austin, Texas

Today

Full-time

USD 123,700.00 per year

Austin, Texas

Today

Full-time

Search all similar jobs