Tampa FL
Job Description:
>> Define and implement evaluation strategies and success metrics for AI agents.
>> Assess task completion, correctness, groundedness, hallucination, tool usage, safety, consistency, and overall user experience of AI-driven systems.
>> Build comprehensive test scenarios, including edge and adversarial cases.
>> Develop automated evaluation pipelines leveraging deterministic checks, LLM-as-a-Judge techniques, and custom evaluators.
>> Conduct detailed examination of agent traces, prompts, retrievals, tool calls, and outputs to identify failure modes and performance bottlenecks.
>> Perform regression testing spanning changes in models, prompts, tools, and workflows.
>> Establish quality thresholds, monitoring dashboards, and release gates.
>> Collaborate with AI engineering, product, business, and risk teams for continuous improvement.
Requirements:
>> Understanding of GenAI, LLMs, RAG, prompts, tool calling, and AI Agents.
>> Proficiency in Python and experience working with APIs, JSON, datasets, and evaluation frameworks.
>> Ability to translate business expectations into measurable evaluation criteria.
>> Experience with AI evaluation observability tools such as LangSmith, Langfuse, Ragas, DeepEval, MLflow or similar is preferred.
>> Strong analytical, problem-solving, and experimentation mindset.