Lead Distributed Systems Engineer - Services Special Projects

Cupertino, CA, US • Posted 17 hours ago • Updated 4 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

✨ Finding the perfect fit...

Job Details

Skills

  • Real-time
  • Analytics
  • Shipping
  • Computer Science
  • Software Development
  • Customer Facing
  • Web Services
  • API
  • Authentication
  • Authorization
  • High Availability
  • Concurrent Computing
  • Multithreading
  • Data Structure
  • Java
  • C++
  • Object-Oriented Programming
  • Systems Analysis/design
  • Spring Framework
  • Performance Tuning
  • Thread
  • Stacks Blockchain
  • JUnit
  • Mockito
  • Testing
  • Gradle
  • Semantics
  • NoSQL
  • Database
  • Network
  • RPC
  • Continuous Delivery
  • Jenkins
  • GitHub
  • GitLab
  • Continuous Integration
  • Automated Testing
  • Amazon Web Services
  • Google Cloud Platform
  • Google Cloud
  • Cloud Computing
  • Docker
  • Kubernetes
  • Streaming
  • Apache Kafka
  • Data Storage
  • Apache Cassandra
  • Caching
  • Redis
  • Apache Avro
  • Management
  • Code Review
  • Threat Modeling
  • Presentations
  • Leadership
  • Collaboration
  • Communication
  • Agile
  • Algorithms
  • Optimization
  • Artificial Intelligence
  • Machine Learning (ML)
  • Cryptography
  • TLS
  • X.509
  • Identity Management
  • OAuth
  • OIDC
  • SAML
  • Lifecycle Management

Summary

Apple's Special Services Team is seeking a Lead (Principal) Distributed Systems Engineer to design and build massively scalable, highly available services that power experiences for Apple customers both now and in the future.

Description

In this Lead role, you will build and operate high-throughput, low-latency backend services that ingest, process, and serve data at scale across a range of mission-critical workloads - from real-time transactions to analytics and content delivery. You'll also drive the evolution of a multi-tenant platform, including AI/ML-powered services, by shipping new capabilities, scaling what exists, and applying distributed-systems best practices from design through production.

Minimum Qualifications

Master's degree in Computer Science or a related field

15+ years of professional software development experience building scalable, distributed systems in production, with at least 5 in a Principal or Sr. Staff Level role.

Experience building, authoring, and operating large-scale, multi-tiered distributed systems and customer-facing web services: including API design, authentication, authorization, scaling for high availability, concurrency, and reliability.

Strong understanding of concurrency and multi-threaded programming, fundamental data structures, and efficient algorithm design

Strong proficiency in Java; working knowledge of a second systems language (Go, C++) is a plus. Solid OO analysis and design skills.

Strong proficiency in application frameworks (Spring boot)

Hands on experience with JVM performance tuning and profiling - selection/tuning, JFR, async-profiler, heap/thread-dump analysis for low-latency services.

Hands-on with async and reactive JVM stacks: Netty, Project Reactor, RxJava, Vert.x, or Micronaut/Quarkus.

Rigorous testing discipline with JUnit 5, Mockito, AssertJ, Testcontainers, and contract testing.

Hands on experience with build and dependency management with Gradle

Deep understanding of transactional consistency models - ACID semantics, with deep knowledge of tradeoffs between relational and NoSQL database technologies

Expertise with synchronous and asynchronous network I/O and RPC frameworks (gRPC)

Experience building and maintaining CI/CD pipelines (e.g., Jenkins, GitHub Actions, GitLab CI, or similar) for automated testing, build, and deployment of production services.

Experience with AWS or Google Cloud Platform and cloud-native tooling (Docker, Kubernetes) in the context of deploying scalable production grade services.

Experience with event streaming and queueing systems (specifically Kafka) and stream processing frameworks and high-throughput, append-only write paths for durable, queryable historical records.

Hands on Experience of leveraging data storage (Iceberg, Cassandra) and caching technologies (Redis) in Production services

Hands on experience with Serialization/schema tooling: Jackson, Protobuf, Avro, and Schema Registry

Experience identifying, triaging, and remediating security vulnerabilities in production services (dependency management, secure code review, threat modeling).

Full life-cycle development experience for a consumer product, from concept through deployment

Proven history of presenting technical and business concepts to Executive Leadership

Preferred Qualifications

Self-motivated, with strong collaboration and communication skills, and experience in a fast-paced, agile environment

Experience with machine learning systems, ML frameworks, libraries and algorithms

Familiarity with deployment and optimization of Large scale Production grade AI Services that require GPUs in the path of the transaction.

Hands-on experience deploying, serving, and optimizing LLMs or ML models directly in the transaction/request path

Experience with security and cryptography (e.g., TLS, X.509 certificates) identity and access management protocols (OAuth2/OIDC/SAML), and secure token/session lifecycle management.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90733111
  • Position Id: 205aeb45692b806e18f71722532bc821
  • Posted 17 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Cupertino, California

Today

Full-time

Mountain View, California

Today

Full-time

USD 150,000.00 - 226,000.00 per year

Cupertino, California

Today

Full-time

San Jose, California

Today

Full-time

USD 184,000.00 - 208,000.00 per year

Search all similar jobs