Lead Generative AI Evaluation Engineer- 7+ yrs- New York, United States- Onsite

New York, NY, US • Posted 3 hours ago • Updated 3 hours ago
Contract Corp To Corp
Contract Independent
Contract W2
6 Months
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Generative AI agents deployed across our Cloud GTM space (e.g.
  • deep research
  • conversational analytics
  • pipeline forecasting).In this role
  • you will transition our AI evaluations from manual human reviews to high-throughput automated pipelines backed by LLM-as-a-Judge and targeted human audits. You will work at the intersection of enterprise software engineering
  • distributed computing
  • and cutting-edge LLM evaluation frameworks to ensure data integrity
  • eliminate regressions
  • and scale our GTM agent capabilities.Key ResponsibilitiesFoundations & Local Sandbox DevelopmentBuild robust golden datasets extracted from UAT logs and utilize frontier models to synthetically generate variations (typos
  • phrasing
  • syntax) for robust testing.Production Pipeline Automation & System ArchitectureArchitect and deploy language-agnostic RPC endpoints to systematically invoke GTM agents within the ecosystem.Build resilient production pipelines supporting parallel inference execution across 1
  • 000+ trajectory datasets in under 15 minutes.Stand up a centralized Model Context Protocol (MCP) logging server to capture raw prompts
  • tool trajectories
  • SQL queries
  • and token costs using strict JSON schemas.Implement asynchronous message queues (Pub/Sub) for rate-limiting/backpressure
  • along with retry policies for network and generation failures.Skill Benchmarking & Trajectory ValidationAuthor eval test suites to isolate specific agent competencies (e.g.
  • CRM writes
  • SQL analytics).Inject sandboxed mocks to validate tool-calling logic without producing live CRM side effects or executing heavy database reads.Stream execution logs for baseline delta analysis and capability scoring.Enterprise Analytics
  • Governance & Security (Phase 4)Build low-latency Hydra ETL pipelines to stream structured evaluation JSON records into data warehouses and construct Plx analytics dashboards

Summary

Job Description :

About the RoleWe are seeking a Lead AI/ML Engineer to architect, build, and operationalize a full-lifecycle, automated evaluation harness for non-deterministic Generative AI agents deployed across our Cloud GTM space (e.g., deep research, conversational analytics, pipeline forecasting).In this role, you will transition our AI evaluations from manual human reviews to high-throughput automated pipelines backed by LLM-as-a-Judge and targeted human audits. You will work at the intersection of enterprise software engineering, distributed computing, and cutting-edge LLM evaluation frameworks to ensure data integrity, eliminate regressions, and scale our GTM agent capabilities.Key ResponsibilitiesFoundations & Local Sandbox DevelopmentBuild robust golden datasets extracted from UAT logs and utilize frontier models to synthetically generate variations (typos, phrasing, syntax) for robust testing.Production Pipeline Automation & System ArchitectureArchitect and deploy language-agnostic RPC endpoints to systematically invoke GTM agents within the ecosystem.Build resilient production pipelines supporting parallel inference execution across 1,000+ trajectory datasets in under 15 minutes.Stand up a centralized Model Context Protocol (MCP) logging server to capture raw prompts, tool trajectories, SQL queries, and token costs using strict JSON schemas.Implement asynchronous message queues (Pub/Sub) for rate-limiting/backpressure, along with retry policies for network and generation failures.Skill Benchmarking & Trajectory ValidationAuthor eval test suites to isolate specific agent competencies (e.g., CRM writes, SQL analytics).Inject sandboxed mocks to validate tool-calling logic without producing live CRM side effects or executing heavy database reads.Stream execution logs for baseline delta analysis and capability scoring.Enterprise Analytics, Governance & Security (Phase 4)Build low-latency Hydra ETL pipelines to stream structured evaluation JSON records into data warehouses and construct Plx analytics dashboards

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91143549
  • Position Id: 5014-31204-1790785891
  • Posted 3 hours ago

Company Info

About iMedhas Consulting Services

Welcome to iMedhas Consulting Services. We are an IT consulting and services enterprises with precision expertise in Digital Transformations, Big data and Analytics. Through our expert team, we provide greater adaptabilities to disseminate with new technologies to increase the overall productivity, operational efficiency, and productivity of a particular business process of an organization.

The goal is the achievement of higher customer satisfaction and providing lucrative returns on investment to the business. On the other hand, we deliver content management expertise, ERP (SAP) systems integration, EAI, and Information Management services.

About_Company_One
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

New York, New York

•

Today

Easy Apply

Contract, Third Party

Depends on Experience

New York, New York

•

2d ago

Easy Apply

Third Party, Contract

Depends on Experience

Search all similar jobs