Senior Data Scientist AI Evaluation_Hybrid@Johnston,RI

Hybrid in Johnston, RI, US • Posted 55 minutes ago • Updated 55 minutes ago
Contract W2
Contract Independent
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Artificial Intelligence
  • Generative Artificial Intelligence (AI)
  • Evaluation
  • Machine Learning (ML)
  • NumPy
  • Financial Services
  • Interactive Voice Response
  • Python
  • Data Science

Summary

Role: Senior Data Scientist Generative AI / Conversational AI Evaluation
Location: Johnston, RI or Westwood, MA

Hybrid/Onsite preferred, though location flexibility may be considered

Domain: Cards, Banking, FSI

Description:

We are seeking an experienced Data Scientist to support the development and evaluation of AI-powered fraud self-service voice agents and conversational AI systems. The primary responsibility is not model deployment or engineering implementation, but designing evaluation frameworks, measuring system performance, identifying failure patterns, conducting root-cause analysis, and optimizing model behavior through data-driven experimentation.

Key Responsibilities

  • Design and execute evaluation frameworks for LLM, RAG, and multi-turn conversational AI systems.
  • Develop metrics to assess customer intent recognition, conversation quality, guardrail effectiveness, and business outcomes.
  • Analyze voice-agent interactions and identify areas of failure, drift, and performance degradation.
  • Perform prompt tuning and experimentation to improve model accuracy and reliability.
  • Conduct root-cause analysis of conversational failures and recommend remediation strategies.
  • Measure performance across different model configurations, prompts, and guardrail implementations.
  • Partner with AI Engineering and Product teams to validate solutions before production deployment.
  • Build dashboards and reports that communicate model effectiveness and operational impact.
  • Support fraud-related customer service use cases, including intent detection and multi-turn conversation flows.

Success Criteria

  • Develop reliable evaluation methodologies for conversational AI systems.
  • Quantify the effectiveness of fraud self-service voice agents.
  • Optimize prompts, retrieval strategies, and guardrails using empirical evidence.
  • Deliver actionable insights that improve customer experience and model performance.
  • Establish measurable KPIs for intent detection and multi-turn conversation success.

Mandatory Skills:

  • Strong background in Data Science, Machine Learning, Generative AI, or a related quantitative field.
  • Hands-on experience evaluating LLM, RAG, Agentic AI, or Conversational AI solutions.
  • Deep understanding of model evaluation techniques and metrics, including:
    • Precision@K
    • Recall@K
    • Mean Reciprocal Rank (MRR)
    • F1 Score
    • Retrieval and generation quality assessment
  • Experience performing experimentation, statistical analysis, and performance benchmarking.
  • Strong Python programming skills.
  • Experience with machine learning libraries and frameworks such as Scikit-learn, XGBoost, Pandas, NumPy, and related tools.
  • Ability to communicate technical findings succinctly to highly technical stakeholders.

Desired Skills:

  • Experience with:
    • Generative AI and LLM ecosystems
    • Multi-agent systems
    • RAG/Agentic RAG architectures
    • Amazon Bedrock
    • AWS SageMaker
    • Databricks
    • MLflow
    • LangSmith
    • Weights & Biases
  • Knowledge of conversational AI, IVR systems, digital assistants, and voice agents.
  • Experience in financial services, fraud detection, or customer service automation
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10500016
  • Position Id: 9086244
  • Posted 55 minutes ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Westwood, Massachusetts

Today

Easy Apply

Contract

Depends on Experience

Johnston, Rhode Island

Today

Easy Apply

Full-time

Depends on Experience

Worcester, Massachusetts

Today

Easy Apply

Contract

$80 - $90

Boston, Massachusetts

30+d ago

Easy Apply

Third Party, Contract

$60,000 - $80,000

Search all similar jobs