100% Remote - QA Automation OR Data Scientist with AI Exp.

Remote • Posted 1 hour ago • Updated 58 minutes ago
Contract W2
6 Months
Remote
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • QA
  • Automation
  • Generative AI
  • Data Scientist
  • LLM

Summary

Job Title: Senior QA Automation Engineer / Data Scientist (Generative AI)

Location: 100% Remote

Interview Process: Video (2 Rounds)

Job Description:

We arelooking for a Data Scientist or QA Automation Engineer with expertise and/or experience in LLM (Large Language Model) evaluation to support quality assurance initiatives for Generative AI solutions. This role focuses on evaluating AI-generated outputs, defining quality metrics, and developing practical approaches for testing non-deterministic systems. Candidates from a variety of backgrounds - software testing, automation, data science, or machine learning - are encouraged to apply. The most important qualification for this role is the ability to research and develop innovative methods for evaluating LLM-based systems, where traditional pass/fail testing approaches may not be sufficient. Experience with DeepEval or similar evaluation frameworks is a plus, but not required.

Roles & Responsibilities

LLM Evaluation & Quality Assurance

Design and execute testing strategies for Generative AI and LLM-powered applications.

Define quality standards, evaluation criteria, and success metrics for AI-generated outputs.

Research and develop innovative approaches for evaluating non-deterministic AI systems.

Create test datasets, benchmark scenarios, and repeatable evaluation methodologies.

Identify and analyze quality issues including hallucinations, inconsistencies, inaccuracies, and reliability concerns.

Develop processes to measure and track AI quality over time.

Automation & Analysis

Build and maintain automated testing and evaluation solutions.

Analyze evaluation results and provide recommendations for quality improvement.

Support integration of AI quality validation into existing development and testing processes.

Develop reports, dashboards, and metrics that communicate AI quality and performance trends.

Collaboration & Innovation

Collaborate with product owners, engineers, data scientists, and business stakeholders to define quality expectations.

Support continuous improvement efforts through experimentation and data-driven analysis.

Stay current with emerging AI evaluation techniques, tools, and industry best practices.

Help establish repeatable quality assurance practices for AI-enabled solutions.

Required Qualifications

Must-Have Skills

Background as a Data Scientist, QA Automation Engineer, or similar technical role, with a degree in a related field or equivalent practical experience.

Expertise and/or hands-on experience with LLM (Large Language Model) evaluation.

Demonstrated ability to research and develop innovative methods for evaluating LLM-based, non-

deterministic systems.

Working knowledge of Python or another programming language commonly used for testing and data

analysis.

Strong analytical, research, and problem-solving abilities, with comfort operating in an emerging technology area where best practices are still evolving.

Good written and verbal communication skills.

Preferred Qualifications (Nice to Have)

Experience with DeepEval or similar LLM evaluation frameworks - a plus, not required.

Broader exposure to Generative AI, Machine Learning, or Natural Language Processing.

Experience evaluating AI-powered applications such as chatbots, copilots, or intelligent assistants.

Familiarity with prompt engineering concepts or major AI platforms (e.g., OpenAI, Anthropic, Google Gemini).

Experience in Financial Services or another regulated industry.

Success Profile

The ideal candidate is naturally curious, analytical, and comfortable working in emerging technology areas where best practices are still evolving. They can evaluate complex problems objectively, rapidly learn new concepts, and create practical approaches for measuring AI quality and performance. Most importantly, they possess the ability to research and develop innovative methods for LLM evaluation and translate those methods into scalable and repeatable testing practices.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91089145
  • Position Id: 2026-13
  • Posted 1 hour ago

Company Info

About SDH Systems

SDH Systems is a team of creatives with experience in design, content, vision and mission. Using these skills we create creative assets.
We always ensure that our software solutions help your business/organization to enhance their productivity by providing you the unmatched developement solutions tailored to suit your needs. Most of our team has on the job experience in more than one medium, so that we can integrate these to build a creative identity.

About_Company_OneAbout_Company_Two
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs