AI Quality Engineer (need only local- Inperson Interview Required)

Bellevue, WA, US • Posted 6 hours ago • Updated 5 minutes ago
Contract Corp To Corp
Contract W2
Contract Independent
12 Months
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Machine Learning
  • Test Planning
  • Artificial Intelligence
  • Regression Testing
  • API Management
  • Data Quality
  • Metrics
  • Data Validation
  • Data Pipelines
  • SQL Databases
  • quality management
  • retrieval-augmented generation
  • large language models
  • Reliability
  • Safety Principles
  • Python (Programming Language)
  • Scalability
  • Acceptance Testing
  • OpenAI
  • Programme Evaluations
  • BLEU Score
  • Hallucination Detection
  • Agentic-AI
  • AI Platforms
  • Chatbots
  • DeepEval
  • Evaluation of Large Language Models
  • LangSmith
  • Low Latency
  • Microsoft Copilot
  • Promptfoo
  • RAGAS (Retrieval Augmented Generation Assessment)

Summary

QA Engineer for AI Products/Solutions- Hybrid
Location Bellevue, WA (In Person Interview Required)

Long Term

Job Description

  • Design and execute test plans for AI/ML-driven features, including model outputs, prompts, and integrated application behaviour
  • Build and maintain automated test suites covering functional, regression, integration, and API testing
  • Evaluate model outputs for accuracy, consistency, bias, hallucination, and edge-case failures
  • Develop evaluation frameworks and golden datasets/test cases to benchmark model performance over time
  • Test prompt engineering changes, model version upgrades, and fine-tuning outputs for regressions
  • Perform adversarial and red-team style testing to surface safety, security, and robustness issues
  • Validate data pipelines feeding into AI models (data quality, schema, drift detection)
  • Collaborate with data scientists/ML engineers to define acceptance criteria and quality metrics for models
  • Test latency, scalability, and reliability of AI services under load
  • Contribute to CI/CD pipelines, integrating automated and model-evaluation tests
  • Hands-on experience testing LLM-based products (chatbots, copilots, RAG systems, AI agents) - designing test cases for non-deterministic, generative outputs
  • Practical experience with AI/LLM evaluation frameworks (e.g., Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, TruLens) - building eval suites, scoring rubrics, and golden datasets
  • Working knowledge of eval metrics for generative AI: hallucination rate, faithfulness/groundedness, relevance, answer correctness, toxicity/bias scoring, BLEU/ROUGE/semantic similarity where applicable
  • Experience with prompt regression testing - validating prompt changes and model/version upgrades against baseline eval sets
  • Strong proficiency in Python for writing eval scripts, test harnesses, and data validation logic
  • Familiarity with SQL and data validation techniques

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90719156
  • Position Id: 2026-12197
  • Posted 6 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Bellevue, Washington

•

Today

Easy Apply

Contract

Depends on Experience

Seattle, Washington

•

6d ago

Easy Apply

Contract

$30 - $37.5

Redmond, Washington

•

Today

Full-time

USD 102,100.00 - 202,200.00 per year

Seattle, Washington

•

Today

Full-time

USD 184,030.00 - 275,400.00 per year

Search all similar jobs