Software Engineer —AI Evaluation & Automation
Remote
12 Months contract
Role Overview
Help build and scale the tooling we use to measure how well AI-powered software
development tools actually perform. You'll develop evaluation harnesses, automate
benchmark runs, and help make sure the results we produce are reproducible and hold up
to scrutiny. This is an engineering role, but a lot of the work is about getting the
measurement right, not just automating it.
Key Responsibilities
• Build and integrate evaluation harnesses and automation for software development use
cases, including turning real engineering artifacts like merged pull requests into
repeatable benchmark tasks.
• Build versioned, repeatable processes to evaluate AI tools, models, and harnesses,
with reproducible run environments (pinned dependencies, containerized runs,
isolated worktrees) so results stay comparable over time.
• Validate and calibrate evaluation approaches against human judgment, so scores are
consistent and correct rather than just repeatable.
• Support execution-based benchmarking across quality, productivity, and eJiciency
measures, including cost and latency.
• Analyze results across repeated runs, looking at variance, failure patterns, and cost per
outcome, and find ways to make the workflows more reliable and more automated.
• Work with engineering and data teams to improve the tooling, and document how the
evaluations work and what they found for both technical and leadership audiences.
Required Skills & Experience
• Strong software engineering background, with real experience building automation,
developer tooling, or test and validation systems.
• Proficient in at least one general-purpose language such as Python, Java, or JavaScript
— the specific language background is flexible.
• Solid working knowledge of Git, including how branches, history, and working trees
behave, and of containerization with Docker.
• Experience with APIs, development environments, CI/CD pipelines, and typical
engineering workflows.
• Understanding of how AI, LLM, or agent evaluation works and where it goes wrong, such
as why a judge can be consistent but still wrong, why a single run can mislead, and how
benchmark contamination happens.
• Able to troubleshoot technical problems, think clearly about whether a measurement is
valid, and analyze results carefully.
• Hands-on experience using AI coding tools and agentic harnesses such as Claude
Code, Devin, or Cursor, and command of the best practices for working with them effectively.
Preferred Experience
• Experience designing benchmarks or evaluations for software systems, especially
execution-based grading that verifies against tests.
• Familiarity with LLM-as-judge or agent-as-judge approaches, and how to check them
against human raters.
• Experience with build-system-aware test selection, such as Bazel or mapping changed
files to the tests that cover them.
• Experience building reproducible test environments and managing versioned
evaluation datasets.
• Comfortable writing up methodology and results for engineering leadership.