Machine Learning Engineer, Foundation Model Services

Santa Clara, CA, US • Posted 30+ days ago • Updated 36 minutes ago
Full Time
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Music
  • Drawing
  • Real-time
  • Build Tools
  • Computer Hardware
  • Use Cases
  • Shipping
  • Python
  • Cloud Computing
  • Amazon Web Services
  • Microsoft Azure
  • Google Cloud Platform
  • Google Cloud
  • Kubernetes
  • Docker
  • Machine Learning (ML)
  • Natural Language Processing
  • Information Retrieval
  • Statistics
  • Communication
  • Collaboration
  • Research
  • Computer Science

Summary

Do you think differently? Are you eager to break the status quo, bold and ambitious, unafraid to take risks, and passionate about building best-in-class technology? If so, there's no better place to do it than Apple. The Foundation Model Services team builds the frameworks, services, and tools that run Apple's largest foundation models in production. Our infrastructure powers intelligent experiences across products people use every day - Search, Music, TV, the App Store, Messages, Photos, Spotlight, Safari, Siri, and more - serving millions of queries at incredibly low latency while drawing every ounce of performance from our hardware. Join us and you'll help bring intelligence to billions of users around the world, working on optimizing and serving large language, vision, and speech models at Apple's scale.

Description

Work closely with product teams to build production-grade solutions that launch models serving customers in real time.

Partner with foundation model researchers to prototype and develop inference for cutting-edge model architectures, and build tools that help us understand and remove performance bottlenecks across different hardware and use cases.

Write high-quality code, learn quickly in a fast-moving field, and grow your impact as you take on larger pieces of the system.

Minimum Qualifications

2+ years of industry experience building and shipping production software and/or machine learning systems.

Proficiency in a modern programming language such as Go or Python.

Experience deploying and operating services on a cloud platform (AWS, Azure, Google Cloud Platform, or equivalent) using containers and Kubernetes/Docker.

5 year+ industry experience in ML technologies (LLMs, Machine Learning, NLP, Information Retrieval, Statistics).

Experience building or operating high-throughput, low-latency services.

Strong communication and collaboration skills, with the ability to partner across research and product teams.

Bachelor's degree or higher in Computer Science or related technical field.

Preferred Qualifications

Familiarity with Nvidia TensorRT-LLM, vLLLM, DeepSpeed, Nvidia Triton Server etc.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90733111
  • Position Id: 6c5548ab174ad3cc936c3a1dd53c4d5c
  • Posted 30+ days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Santa Clara, California

•

Today

Full-time

USD 158,900.00 - 178,100.00 per year

Sunnyvale, California

•

Today

Full-time

USD 169,000.00 - 338,000.00 per year

San Jose, California

•

Today

Full-time

USD 151,800.00 - 265,350.00 per year

Palo Alto, California

•

Today

Full-time

USD 209,000.00 - 313,000.00 per year

Search all similar jobs