Senior Applied Scientist, Multilingual AI Evaluation

Washington, WA, US • Posted 7 days ago • Updated 8 minutes ago
Full Time
On-site
Fitment

Dice Job Match Score™

🤯 Applying directly to the forehead...

Job Details

Skills

  • Science
  • Computer Science
  • Research
  • Python
  • Artificial Intelligence
  • Language Models
  • Shipping
  • Collaboration
  • Communication
  • Linguistics
  • Computational Linguistics
  • FOCUS
  • Modeling
  • Publications
  • Natural Language Processing
  • Multilingual
  • Machine Learning (ML)
  • PyTorch
  • JAX
  • Internationalization And Localization
  • Workflow
  • Evaluation
  • Fluency
  • English

Summary

AI systems are only as trustworthy as the methods used to evaluate them. At Apple, where AI powers experiences for billions of people around the world, getting evaluation right is not a support function-it is a foundational science. Our team, part of Apple Services Engineering, is building that scientific foundation: rigorous, scalable evaluation methodology for LLMs, agentic systems, and human-AI interaction.

We're looking for a senior applied scientist to make the evaluation tooling we build work across every language and culture Apple serves. This is a role for someone who is fluent in both modern AI and the science of language, and who can set direction and drive initiatives independently, not just execute them. You'll do this on a deeply interdisciplinary team working alongside ML researchers, measurement scientists, and platform engineers.

Description

In this role, you'll help ensure Apple's AI features work well across languages and cultures. Your goal is to make our evaluation tooling multilingual from the start so that engineers building AI features can design, test, and ship across the world from day one. It's a broad applied science role: you'll shape how Apple evaluates AI wherever the hardest questions are, and you'll have the opportunity to publish novel work.

The scientific challenge is real. How do we ensure we consistently evaluate AI features across different grammar, script, or cultural norms and how do we do this at scale? You'll bring linguistic judgment to questions like these and, working with measurement scientists and ML researchers, turn it into validated methodology that holds across dozens of languages.

This is a hands-on role. You'll design and implement your own methods in Python, working closely with research and engineering partners, while staying focused on the science of getting evaluation right.

Minimum Qualifications

MS in Linguistics, Computational Linguistics, NLP, Computer Science, or a related field - or equivalent research/work experience.

Deep expertise in linguistics, with working fluency in the structure of multiple languages beyond English.

Strong proficiency in Python.

Solid understanding of LLMs and AI evaluation fundamentals, including how language models process and generate across languages.

Demonstrated experience shipping or evaluating features across multiple languages or locales.

Experience designing benchmarks, datasets, or human evaluation protocols, with attention to statistical rigor and reproducibility.

Ability to drive initiatives independently and collaborate across a cross-functional, interdisciplinary team.

Strong written and verbal communication skills.

Preferred Qualifications

PhD in Linguistics, Computational Linguistics, or NLP with a focus on multilingual or cross-lingual modeling.

Publications in NLP, multilingual evaluation, or evaluation methodology.

Hands-on experience with modern ML frameworks (PyTorch, JAX) and with fine-tuning or evaluating LLMs.

Experience with low-resource languages, dialectal variation, or sociolinguistics.

Familiarity with localization/internationalization workflows and quality assessment.

Experience with LLM-as-judge approaches, rubric design, or bias and fairness evaluation across languages.

Fluency or professional proficiency in one or more languages in addition to English.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90733111
  • Position Id: 8ca18cecc68c882ff1a676443e82c3a2
  • Posted 7 days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Washington

Today

Full-time

Remote

Today

Full-time

USD 180,000.00 - 230,000.00 per year

Remote

Today

Full-time

Remote

14d ago

Easy Apply

Full-time

35 - 40

Search all similar jobs