The Apple Intelligence Platform Experience Validation team builds the tooling and automation that keeps Apple Intelligence features high-quality before they ship. We are looking for a Senior SDET to lead the design and implementation of automated model evaluation: standing up LLM-as-a-judge in existing and new pipelines, and building the infrastructure that catches model regressions before they reach human evaluation or the live on population.
This is a hands-on, senior individual-contributor role. You will own eval automation as a discipline across the team, partnering with modeling, framework, and infrastructure teams to make model quality a first-class, continuously measured signal.
Description
You will build and maintain model level, component or end-to-end evaluation coverage for the generative features our team validates. Your job is to leverage LLM judge scoring output quality in automation, ensuring reliable, repeatable eval jobs that run that produce actionable signal.
The kinds of problems you will work on include:
* Image / visual generation: validating model output and its associated classification metadata, and detecting quality or behavior regressions across model updates.
* Natural-language generation: evaluating whether generated artifacts and responses match user intent, moving at-desk LLM judges into a scalable and repeatable automation environment.
* Correctness beyond string matching: replacing exact-match checks for open-ended or factual responses with an LLM-as-judge stage integrated into the pipeline.
* Generated insights and summaries: assessing whether model-generated content is sensible and good enough to surface to users.
You will decide when a component-level check (an API or CLI that exercises the model against its framework) is sufficient and when a full end-to-end user flow is required, and you will build the tooling for both.
Minimum Qualifications
BS in Computer Science, Mathematics, or a related field (or equivalent practical experience)
Three years of relevant industry experience in test automation, software development, or related areas.
Preferred Qualifications
Strong practical knowledge of Python, including data-pipeline fluency (JSON/YAML, REST APIs).
Hands-on experience with LLM-as-a-judge evaluation and rubric design, or a strong demonstrated ability to ramp into it quickly.
Strong software engineering fundamentals - able to define atomic, composable components and build maintainable pipelines and tooling, not just scripts.
Strong debugging and triage skills; able to separate genuine regressions from infrastructure or rubric noise.
Strong knowledge of the software development lifecycle, testing methodologies, and QA processes.
Excellent written and verbal communication; able to document clearly and describe quality signal to modeling and leadership audiences.
Ability to lead work across varying priorities and partner multi-functionally with modeling, framework, and infrastructure teams.
Experience building on-device tooling and device/model eval infrastructure.
Experience integrating with CI/CD and job orchestration systems, and comfort deploying tooling as reusable libraries.
Familiarity with generative model behavior - image generation, NLP, or LLM output evaluation.
Experience curating and reasoning about large datasets; comfort manually inspecting data (Jupyter or similar) to build intuition and drive next steps.
Awareness of dataset bias and fairness considerations in evaluation.
Experience with database/query tooling (e.g., SQL) and dashboards/visualization for reporting quality trends.
Experience with Xcode is a bonus.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 90733111
- Position Id: 6342395f70c33cfbd0d09c03c333f0f6
- Posted 1 hour ago