Remote or Hybrid in New York, New York
•
Today
Role: AI Evaluations Engineer (Senior) Location: US (remote) Duration: 12 Months + Key Responsibilities Design and operate end-to-end evaluation frameworks for agentic AI systems (offline regression, online monitoring, A/B testing) Define quality rubrics, scoring architectures, and evaluation standards across multi-agent workflows Build LLM-as-judge and Agent-as-judge pipelines assessing trajectory, tool-call accuracy, groundedness, and safety Own and deliver evaluation infrastructure end-to-en
Easy Apply
Contract


