Observability & Evaluation Engineer - NTTJP00113656

Charlotte, NC, US • Posted 1 day ago • Updated 26 minutes ago
Contract W2
Contract Corp To Corp
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Agile
  • PYTHON
  • AI/ML
  • LLM

Summary

TECHNOGEN, Inc. is a Proven Leader in providing full IT Services, Software Development and Solutions for 15 years.

TECHNOGEN is a Small & Woman Owned Minority Business with GSA Advantage Certification. We have offices in VA; MD & Offshore development centers in India. We have successfully executed 100+ projects for clients ranging from small business and non-profits to Fortune 50 companies and federal, state and local agencies.


Observability & Evaluation Engineer
Charlotte, NC (Onsite)
Description

Observability & Evaluation Engineer will build telemetry, tracing, dashboards, evaluation suites, alerts, service objectives, runbooks, and readiness evidence for Tachyon agent releases. This role ensures production AI systems can be monitored, evaluated, improved, and supported with clear operational visibility.

Key Responsibilities

  • Implement observability and telemetry for LLM-powered applications, agents, tools, and platform services.
  • Build evaluation suites for agent behavior, prompt quality, response quality, retrieval performance, latency, reliability, and safety signals.
  • Develop dashboards, alerts, traces, metrics, service objectives, and reporting for production readiness.
  • Work with platform engineers and Product Owners to define monitoring requirements and evaluation metrics.
  • Automate evidence collection for release readiness, operational reviews, and governance checkpoints.
  • Create runbooks and support documentation for priority agent releases.
  • Analyze production behavior and recommend improvements to reliability, performance, and quality.

Required Qualifications

  • 7+ years of engineering experience with observability, monitoring, test automation, platform operations, or AI/ML systems.
  • 5+ years of strong hands-on Python experience.
  • 5+ years of Experience with dashboards, metrics, alerts, traces, logs, SLOs, and production monitoring.
  • 5+ years of Understanding of LLM evaluation, prompt evaluation, RAG evaluation, or AI quality assessment approaches.
  • 5+ years of Experience working in Agile engineering teams and production support environments.

Required Skills / Knowledge

  • Python, telemetry, tracing, monitoring, dashboards, alerting, SLOs, evaluation frameworks, test automation, and production operations.
  • Understanding of LLMs, agents, RAG, prompt performance, retrieval quality, latency, and reliability metrics.
  • Experience with observability tools and open telemetry concepts.
  • Preferred Qualifications
  • Experience with GenAI observability, AI evaluation tools, ML monitoring, or platform reliability engineering.
  • Experience in regulated environments with evidence and readiness documentation.
  • Kubernetes, cloud platforms, and CI/CD experience.
  • Expected Outcomes
  • Operational dashboards and evaluation suites for priority agent releases.
  • Clear readiness evidence, alerts, SLOs, and runbooks.
  • Improved quality, reliability, and trust in production Agentic AI systems.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10217412
  • Position Id: 2026-43538
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Charlotte, North Carolina

•

Today

Easy Apply

Contract

Depends on Experience

Charlotte, North Carolina

•

Today

Full-time

Charlotte, North Carolina

•

Today

Full-time

USD 96,800.00 - 145,200.00 per year

Charlotte, North Carolina

•

Today

Full-time

Search all similar jobs