- 5+ years of experience in ML engineering, MLOps, platform engineering, or SRE, including 2+ years working hands-on with LLMs or LLM-powered applications in production.
- Demonstrated experience building evaluation systems for ML or LLM applications: test harnesses, benchmark datasets, automated scoring (including LLM-as-judge approaches), and regression detection.
- Strong software engineering skills in Python (and ideally TypeScript), with a track record of building reliable, well-tested internal platforms and tooling.
- Deep familiarity with CI/CD systems (e.g., GitHub Actions, GitLab CI, Jenkins, Buildkite) and experience embedding automated quality gates into deployment pipelines.
- Experience with observability and monitoring stacks (e.g., OpenTelemetry, Datadog, Grafana/Prometheus) and, ideally, LLM-specific observability tools (e.g., LangSmith, Langfuse, Arize Phoenix, Braintrust, W&B Weave).
- Proven ability to debug complex distributed systems under pressure, including production incident response, root-cause analysis, and blameless postmortems.
- Excellent cross-functional communication: able to translate evaluation results into clear findings and recommendations for both engineers and non-technical stakeholders.
- Comfort with ambiguity and a builder s mindset: this role starts with a blank page and ends with the evaluation platform the whole organization relies on.
- Experience with agentic frameworks and orchestration patterns (e.g., multi-agent systems, tool use, RAG pipelines) and their distinct failure modes.
- Experience with prompt management, model routing, or fine-tuning workflows and evaluating changes across model versions and providers.
- Background in statistics or experimentation (A/B testing, significance testing, sampling strategies for human review).
- Design and build reusable AI agent skills, plugins and maintain internal marketplace infrastructure to extend and scale Data, AIML capabilities across the organization.
- Expertise in causal inference and measurement strategy including causal graphs, ontologies, and knowledge graphs to drive rigorous, decision grade data analysis.
- Experience operating in regulated or high-stakes domains where agent errors carry real business or customer impact.
|