Role: QA Engineer – AI Applications (Fully Remote)
Domain: Healthcare / Medicaid Analytics
Technology Focus: Generative AI, Agentic AI, AWS Bedrock, Databricks
Duration: 3+ months
Role Summary
We are looking for an experienced QA Engineer to validate an enterprise Generative AI and Agentic AI solution built using AWS Bedrock and Databricks. The role will focus on conversational AI, agent workflows, natural-language-to-SQL, KPI assessments, surveillance capabilities, data accuracy, security, and end-to-end AI quality. The candidate should also have a basic understanding of the US healthcare / Medicaid domain so that business scenarios and AI responses can be validated in the correct functional context.
Key Responsibilities
• Define and execute test scenarios for Generative AI and Agentic AI workflows.
• Validate natural-language questions and AI-generated responses against agreed source-of-truth results.
• Test NL-to-SQL generation, query execution, validation, repair, and result interpretation.
• Validate AI agent workflows, tool/function calls, routing, retries, fallback behavior, and error handling.
• Test KPI threshold alerts, automated assessments, root-cause analysis, evidence, and recommended actions.
• Validate surveillance/anomaly-detection workflows against expected historical outcomes.
• Perform functional, integration, regression, API, data, and end-to-end testing.
• Validate AI responses across different roles/personas and entitlement scenarios.
• Test ambiguous questions, unsupported requests, refusal behavior, and negative scenarios.
• Validate PHI/PII handling, role-based access, and security controls.
• Validate data/results against Databricks and governed business definitions.
• Work with healthcare/business SMEs to understand use cases, expected results, business rules, and KPI definitions.
• Support performance/concurrency testing, UAT, defect triage, production-readiness validation, and hypercare.
• Work with AI Engineers and the AI Architect to isolate failures across data/context, prompts, models, agents, and application logic.
Must-Have Skills
• Strong experience in QA/testing of enterprise applications.
• Strong API testing experience.
• Good SQL and data-validation skills.
• Experience with functional, integration, regression, and end-to-end testing.
• Experience testing complex data-driven applications.
• Understanding of Generative AI / LLM concepts.
• Ability to validate AI-generated results against expected business outcomes.
• Experience with test automation tools/frameworks.
• Basic understanding of US healthcare / Medicaid domain concepts and healthcare data.
• Strong defect analysis, troubleshooting, communication, and documentation skills.
Basic Healthcare Domain Knowledge
• Basic understanding of the US healthcare and Medicaid ecosystem and common MMIS concepts.
• Familiarity with common healthcare data domains such as members/beneficiaries, eligibility, providers, claims/encounters, and pharmacy data.
• Basic understanding of healthcare KPIs, measures, business rules, and the importance of validating results against an approved source of truth.
• Awareness of HIPAA, PHI/PII handling, least-privilege access, and role/persona-based data access.
• Ability to work with healthcare SMEs to translate business scenarios into test cases and expected outcomes.
• Pharmacy / PBM / Medicaid claims knowledge is a plus, but deep domain expertise is not required.
Preferred Skills
• Hands-on experience testing Generative AI / LLM applications.
• Experience testing Agentic AI / AI agents.
• AWS Bedrock exposure.
• Databricks / Unity Catalog experience.
• Natural Language-to-SQL testing.
• RAG / grounding validation.
• AI evaluation frameworks and automated scoring.
• Prompt/model regression testing.
• API automation and performance/load testing.
• Healthcare / Medicaid / Pharmacy domain experience.
• PHI/PII and healthcare-security testing experience.
AI-Specific Testing Experience Preferred
• Response accuracy and factual consistency
• SQL/result accuracy
• Hallucination and unsupported-answer handling
• Grounding and context retrieval
• Tool-call execution
• Conversation/follow-up context
• Agent workflow failures and recovery
• Role/persona-based response differences
• Model/prompt regression
• Latency and cost-related behavior