As part of the Siri organization, you will build the systems and tooling that make evaluation a first-class part of how Siri is developed - not an after-the-fact check - spanning human evaluation, real user feedback, reward and alignment signals, and data science rigor across iOS, iPadOS, macOS, watchOS, and visionOS.
This is a rare opportunity to work at the intersection of software engineering and rigorous evaluation science - applying machine learning engineering techniques, from model evaluation to reward modeling, to build the infrastructure that keeps this quality signal trustworthy. What you build will directly shape the direction of one of the world's most widely used assistants.
Description
In this role you'll contribute across several interconnected work streams spanning evaluation quality, reward/alignment signals, and data science. A core part of the job is bringing "evals first" thinking to the team - building tooling and harnesses grounded in real workflows, and designing for observability and reproducibility so quality can be measured clearly and issues caught early. This is a largely unexplored space with few established playbooks, so being self-driven is a must - you'll define your own path as much as execute one. Scope and priorities will evolve, and we're looking for someone who moves fluidly across these areas, bringing strong software engineering fundamentals with enough ML/LLM depth to build and ship AI-facing tooling
Minimum Qualifications
Strong programming skills in one or more compiled languages (Swift, C++, or Objective-C)
Strong Python skills and solid computer science fundamentals, including data structures, algorithms, and clean, testable code
Ability to quickly learn and adapt to evolving technologies and tools, such as GenAI-assisted coding, new ML frameworks, and emerging LLM/agent tooling
Experience with backend/API development and production debugging
Excellent communication and cross-team collaboration skills, with experience working effectively within large, cross-functional organizations
M.S. or B.S. in Computer Science, Machine Learning, or a related field (or equivalent experience)
Preferred Qualifications
Experience evaluating ML, LLM, or agent-based systems, including familiarity with metrics, scoring methodology, trajectory and outcome analysis, and techniques like prompting, RAG, or LLM as judge
Understanding of reinforcement learning and the underlying techniques behind modern LLMs (e.g. transformer architectures, RLHF/RLAIF, fine-tuning, reward modeling ) and frameworks such as PyTorch or Hugging Face, as applied to evaluation and reward signal design
Familiarity with eval-driven development - defining success criteria and test cases from product goals and real user workflows rather than abstract benchmarks
Experience with data science methods applied to quality measurement - defining ground truth, measuring inter-rater agreement (e.g. Cohen's/Fleiss' kappa), and validating automated scorers using basic statistical techniques (e.g. correlation, confidence intervals, hypothesis testing)
Experience with MLOps, deployment, and test/eval environment management - containerization, CI/CD, model versioning, monitoring, cloud platforms (AWS, Google Cloud Platform, or similar), and staging or provisioning environments to produce repeatable, deterministic conditions
Comfort communicating and collaborating effectively across multicultural teams and time zones
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 90733111
- Position Id: eab78a0a7d53b254e706c4cf3b29036d
- Posted 6 hours ago