An enterprise financial services organization Virginia is seeking an AI Test Engineer to join their team and focus on building automated AI testing frameworks, evaluation pipelines, and quality engineering practices for generative AI, LLM, RAG, and AI-enabled developer experience solutions.
About the Opportunity:
Responsibilities:
Design, develop, and maintain AI testing and evaluation frameworks for generative AI, predictive AI, LLM, RAG, and AI-enabled developer experience solutions
Create and execute automated test suites to validate AI model behavior, output quality, accuracy, reliability, consistency, and performance
Develop automated evaluation pipelines integrated with CI/CD workflows to support continuous quality validation
Perform functional, regression, integration, performance, scalability, and scenario-based testing of AI-enabled applications and services
Analyze AI-generated outputs, investigate defects, and document testing findings, metrics, and recommendations for stakeholders
Qualifications:
3+ years of experience in software testing, quality engineering, test automation, AI/ML testing, or a related discipline
Bachelor's degree in Computer Science, Information Systems, Software Engineering, Data Science, or a related technical field
Strong programming and scripting skills in Python
Experience designing and implementing automated test strategies for complex software applications
Hands-on experience building automated test frameworks and testing pipelines
Experience testing APIs, integrations, microservices, or distributed applications
Experience working with CI/CD environments and modern software delivery practices
Working knowledge of machine learning, generative AI, large language models, or AI-powered applications
Experience with data validation, test-result analysis, defect investigation, and root-cause analysis
Knowledge of modern software testing methodologies, quality engineering practices, and software development lifecycles
Knowledge of statistical analysis and evaluation techniques used to validate system performance
Experience with Git or comparable source control systems and collaborative development practices
Desired Skills:
Experience testing generative AI applications, LLM-based solutions, AI agents, or Retrieval-Augmented Generation systems
Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith, or comparable tools
Experience with LLM observability, monitoring, tracing, and evaluation practices
Experience with Azure AI, GitHub Copilot, OpenAI, Hugging Face, LangChain, or comparable AI platforms and frameworks
Familiarity with prompt engineering, benchmark dataset creation, prompt libraries, and AI output validation techniques
Understanding of Responsible AI principles, AI governance, model risk management, and enterprise AI controls
Experience with performance, scalability, reliability, and resilience testing
Experience supporting Developer Experience, developer productivity, software engineering enablement, or internal developer platforms
Experience working within enterprise, financial services, regulated, or highly controlled technology environments
Experience developing quality dashboards, automated reporting, or AI evaluation analytics
Familiarity with model comparison, A/B testing, experimentation, and statistical evaluation methodologies
Strong curiosity about emerging AI technologies and a passion for building scalable quality engineering capabilities