AI Engineer - Algorithm Evaluation & Agentic Systems

Sunnyvale, CA, US • Posted 19 hours ago • Updated 6 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

🛠️ Calibrating flux capacitors...

Job Details

Skills

  • SAFE
  • Systems Design
  • Algorithms
  • Communication
  • Art
  • Training
  • Probability
  • Computer Vision
  • Language Models
  • Modeling
  • Video
  • KPI
  • Generative Artificial Intelligence (AI)
  • Regression Testing
  • Reasoning
  • Management
  • Workflow
  • Failure Analysis
  • Machine Learning (ML)
  • Python
  • Deep Learning
  • PyTorch
  • Evaluation
  • Mentorship
  • Statistics
  • Testing
  • Design Of Experiments
  • Decision-making
  • Artificial Intelligence
  • Benchmarking

Summary

How do we ensure Apple's next-generation AI products are robust, safe, and truly intelligent? Join the DAQ team to help answer that. We are seeking an AI Engineer specializing in algorithm evaluation and agentic systems design for advanced computer vision and video understanding algorithms.

What We Value

Production mindset: correctness, observability and maintainability

Ability to reason about system-level tradeoffs, not just model performance

Ability to balance experimentation speed with engineering rigor

Comfort working in ambiguous problem spaces and defining metrics from first principles

Clear communication of technical findings to both technical and non-technical audiences

Description

Within the DAQ team, our core mission is to evaluate and elevate advanced visual technologies. As a key member of this group, you will lead the benchmarking and integration of state-of-the-art models for image and video understanding. Rather than focusing on core model training, you will apply your deep CV and ML expertise to rigorously test models in applied settings, uncover edge-case failure modes, and architect advanced agentic systems. If you are passionate about AI safety, robust evaluation, and building autonomous multi-modal workflows that bridge experimentation with production, we'd love to hear from you.

Minimum Qualifications

MS and a minimum of 3 years relevant industry experience

3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation

Solid ML Foundation: Deep understanding of core Machine Learning principles, including probability, statistics, data distributions, and model bias/variance. You can apply statistical rigor to ensure evaluation metrics are meaningful and reliable.

Computer Vision Expertise: Deep theoretical and practical understanding of Computer Vision (CV) and Vision-Language Models (VLMs). You must understand how Vision Transformers (ViTs), spatial-temporal modeling, and image/video processing work under the hood to effectively evaluate them.

Advanced Evaluation Skills: Proven track record of defining robust metrics/KPIs and designing rigorous evaluation frameworks for generative AI or foundation models. Deep experience with custom benchmark creation, automated regression testing, LLM/VLM-as-a-judge methodologies, and human-in-the-loop evaluation.

Agentic Systems: Experience building and evaluating LLM/VLM-powered agents, including tool use, multi-step reasoning, planning, and memory management workflows.

Failure Analysis: Strong intuition for probing ML models to discover edge cases, hallucinations, and performance bottlenecks in constrained environments. Be able to translate findings into actionable improvement recommendations.

Engineering Excellence: Strong proficiency in Python and experience with deep learning frameworks (PyTorch) for running inference, extracting embeddings, and building scalable evaluation pipelines.

Preferred Qualifications

Demonstrated ability to lead technical evaluation strategies end-to-end, drive architectural decisions for testing infrastructure, and mentor engineers.

Strong foundation in statistics, including hypothesis testing, confidence intervals, and experimental design

Knowledge of reinforcement learning, planning, or decision-making systems

Experience evaluating multi-modal or multi-agent systems

Prior work on AI reliability, safety, or benchmarking
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90733111
  • Position Id: 3bb7f53d1f9c626fbc2def22a958304b
  • Posted 19 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Cupertino, California

Today

Full-time

Santa Clara, California

Today

Full-time

USD 201,300.00 - 352,300.00 per year

Cupertino, California

Today

Full-time

Cupertino, California

Today

Full-time

Search all similar jobs