Software Engineer, Platform Reliability Engineering, AiDP

Sunnyvale, CA, US • Posted 1 day ago • Updated 3 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Innovation
  • Application Development
  • Analytics
  • Reliability Engineering
  • Optimization
  • Generative Artificial Intelligence (AI)
  • Computer Science
  • Computer Engineering
  • Python
  • Java
  • Scalability
  • Cloud Computing
  • Data Processing
  • Kubernetes
  • DevOps
  • Management
  • Open Source
  • Systems Architecture
  • Collaboration
  • Big Data
  • Apache Spark
  • Apache Flink
  • Machine Learning (ML)
  • Artificial Intelligence
  • Linux
  • Database
  • Communication
  • Articulate
  • IT Management
  • Operating Systems

Summary

AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning - including generative AI - along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.

The Applied Machine Learning team in AI and Data Platform organization is building the foundation for Apple's enterprise-wide machine learning and data capabilities. Our Applied Machine Learning team designs, builds, and operates mission-critical platforms and services spanning ML, GenAI, inference, and big data-enabling teams across the company to harness AI and analytics at scale. We tackle complex technical challenges in reliability, performance, and scalability across a diverse ecosystem of open source and cutting-edge technologies, serving some of Apple's most demanding workloads.

Description

We're seeking an experienced software engineer to join our Platform Reliability Engineering team and drive the design, operation, and optimization of large-scale distributed systems that power our GenAI, ML, and big data platforms. You'll leverage cutting-edge open source technologies in hybrid cloud environments to build resilient infrastructure that enables seamless inference, data processing, and machine learning workloads at scale. In this role, you'll own mission-critical platform components, respond to production incidents, and collaborate across teams to shape the future of our data and AI infrastructure.

Minimum Qualifications

Bachelor's degree in Computer Science, Computer Engineering, or equivalent professional experience

Proficiency in at least one systems programming language (Python, Go, Java, or similar)

Strong expertise in distributed systems architecture, with deep knowledge of reliability, scalability, and containerization principles

Hands-on experience with cloud platforms and data processing infrastructure (Kubernetes, Spark, Flink, Ray, Trino, or equivalent technologies)

Preferred Qualifications

7+ years of experience in SRE, DevOps, or infrastructure engineering, with demonstrated expertise managing distributed systems at scale.

Proficiency in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed environments.

Familiarity with open source codebases; ability to read, understand, and explain complex system implementations

Strong understanding of system architecture and proven ability to collaborate effectively across engineering teams

Hands-on experience with big data technologies (Spark, Flink, Iceberg) and/or ML/AI platforms (Ray, MLflow, model serving infrastructure).

Strong foundational knowledge of Linux, databases, and security principles

Proactive mindset with demonstrated commitment to optimizing reliability and uptime for mission-critical services

Excellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadership

Demonstrated track record of designing and operating systems at scale
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90733111
  • Position Id: c05f171ae377bec6629d19af6cb44186
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

San Jose, California

Today

Full-time

USD 124,000.00 - 271,200.00 per year

Palo Alto, California

Today

Full-time

USD 165,000.00 - 280,000.00 per year

Menlo Park, California

Today

Full-time

USD 236,000.00 - 339,200.00 per year

San Jose, California

Today

Full-time

USD 173,500.00 per year

Search all similar jobs