AI Evaluation Infrastructure & Production Readiness

Fort Worth, TX, US • Posted 6 hours ago • Updated 6 hours ago
Full Time
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • AI Evaluation
  • Infrastructure
  • Production

Summary

Job Title: AI Evaluation Infrastructure & Production Readiness
Lead Work Location: Various Location(Onsite)

Long Term Contract

Must live within 50 miles of OneMain Corporate Office :: Baltimore, MD | Evansville, IN | Charlotte, NC | Fort Worth, TX | Irving, TX | Wilmington, DE

Function: Technology Enterprise AI (Quality & Risk Controls)

Level: Senior Engineer / Staff Engineer

Reports to: Assoc. Director, GenAI Platform

About the Role :

AI quality cannot be proven once at launch it is an ongoing discipline. Because AI systems are probabilistic, quality can shift as models, prompts, content, vendors, retrieval, and user behavior change. This role builds the evaluation discipline and tooling that produces repeatable evidence for whether an AI system is accurate, safe, compliant, grounded, and ready to scale. It is the enterprise's most important AI risk-control role.

What You'll Build :

  • Evaluation harnesses and tooling that can be run repeatedly across AI systems.
  • Golden test sets and scenario libraries covering expected behaviors and edge cases.
  • Regression testing to catch quality changes when models, prompts, or content change.
  • Hallucination testing and grounding/faithfulness testing.
  • Bias and fairness testing to detect discriminatory or disparate-impact outcomes across protected classes, in support of fair-lending obligations (e.g., ECOA / Reg B).
  • Adversarial and red-team testing, including jailbreak, prompt-injection, and harmful-output resistance.
  • Policy-adherence testing against enterprise, compliance, and regulatory requirements.
  • Drift detection and ongoing production (online) monitoring sampling live traffic, canary evaluations, and catching regressions after release.
  • Human and subject-matter-expert evaluation workflows, plus curation, labeling, and versioning of evaluation datasets (including governance of any customer data they contain).
  • Production-readiness gates that AI systems must pass before scaling.
  • Evaluation evidence and reporting that AI governance and model-risk committees rely on to make go/no-go decisions.
  • AI vendor acceptance criteria for third-party agents and solutions.

Why This Role Matters :
Without a rigorous evaluation function, AI scales on the strength of demos, pilots, anecdotes, or vendor claims rather than repeatable evidence. This role creates the evidence system leadership needs to decide, with confidence, whether an AI system is ready. Evaluations apply equally to agents built in-house and to third-party agents delivered by vendors.

Required Qualifications :

  • 7+ years of relevant experience in software quality, data science, machine learning, or a closely related field, including experience leading a technical workstream.
  • Demonstrated experience designing evaluation methods for LLM or ML systems (accuracy, grounding, safety, and policy adherence).
  • Strong grasp of testing methodology: test-set design, regression testing, and metrics that hold up over time.
  • Statistical rigor: experimental design, significance testing, and sample sizing.
  • Experience with bias and fairness evaluation; familiarity with fair-lending concepts (ECOA / Reg B,disparate impact) is a strong plus.
  • Ability to define and enforce production-readiness gates and acceptance criteria.
  • Comfort partnering with risk, compliance, and legal to translate requirements into testable criteria.

Preferred Qualifications :

  • Experience building evaluation harnesses, LLM-as-judge pipelines, or automated eval frameworks.
  • Familiarity with hallucination, faithfulness, and drift-detection techniques.
  • Experience setting vendor acceptance criteria and evaluating third-party AI solutions.
  • Model risk management or model validation background (e.g., SR 11-7-aligned practices).
  • Background in regulated financial services.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91135807
  • Position Id: 9089099
  • Posted 6 hours ago
Contact the job poster
AT

Atulit Tripathi

Recruiter @ IBOTIX US Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Southlake, Texas

Today

Full-time

USD 82,000.00 - 105,000.00 per year

Irving, Texas

Today

Easy Apply

Third Party

Irving, Texas

9d ago

Easy Apply

Third Party, Contract

$72 - $82

Dallas, Texas

9d ago

Easy Apply

Contract

50

Search all similar jobs