AI DevOps/Observability Engineer

Charlotte, NC, US • Posted 6 hours ago • Updated 6 hours ago
Contract Independent
Contract W2
Contract Corp To Corp
12 Months
No Travel Required
Able to Sponsor
On-site
$55 - $60/hr
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • API
  • Amazon Web Services
  • Artificial Intelligence
  • Good Clinical Practice
  • Evaluation
  • Grafana
  • IaaS
  • Kubernetes
  • DevOps
  • Dashboard
  • Machine Learning Operations (ML Ops)
  • Python
  • Microsoft Azure
  • Performance Analysis
  • LangChain
  • Generative Artificial Intelligence (AI)
  • Incident Management
  • LangSmith
  • Docker
  • Management
  • LlamaIndex
  • Reliability Engineering
  • Testing
  • New Relic
  • Analytical Skill
  • Continuous Integration and Development
  • Google Cloud Platform
  • Regression Analysis
  • Continuous Integration
  • Cloud Computing

Summary

Role : AI DevOps/Observability Engineer

Location : Charlotte, NC (Onsite)

Persistent Systems

 

We are seeking a highly skilled AI DevOps/Observability Engineer to join our production operations team. In this role, you will be responsible for the reliability, performance, and operational readiness of our priority Generative AI and LLM agent releases. You will bridge the gap between AI development and production operations by implementing robust telemetry, automated evaluation pipelines, and comprehensive monitoring systems. Your work will ensure our intelligent agents are stable, accurate, efficient, and scale seamlessly in production environments.

 

Key Responsibilities

  • Observability & Telemetry: Implement and maintain deep tracing, logging, and telemetry solutions specifically tailored for LLM application architectures and multi-agent workflows.
  • Monitoring & Insights: Design, build, and maintain production dashboards that track systemic health, infrastructure metrics, and specialized AI performance indicators.
  • Operational Readiness: Establish actionable alerting systems, define Service Level Objectives (SLOs), curate runtime runbooks, and provide concrete engineering evidence for production readiness.
  • Evaluation & Testing Pipelines: Develop automated continuous evaluation suites to assess LLM agent behavior, safety, and output quality prior to and during deployment.
  • Performance Analysis: Analyze prompt efficiency, token usage, latency, and overall model performance to optimize cost, speed, and accuracy.
  • Production Operations: Support the deployment pipeline, participate in incident management, and continually improve the resilience of our live AI services.

 

Required Skills and Qualifications

Core Technical Skills

  • LLM & Agent Evaluation: Experience with framework-based evaluation tools (e.g., Ragas, DeepEval, TruLens) to measure hallucination, faithfulness, and relevancy.
  • Tracing & Telemetry: Proficiency with LLM-specific tracing tools (e.g., LangSmith, LangFuse, Phoenix, Arize) and open standards like OpenTelemetry.
  • Metrics & Dashboards: Hands-on experience building monitoring views in platforms like Datadog, PrometheGrafana, New Relic, or cloud-native suites.
  • Production Alerting & SLOs: Proven ability to define meaningful Service Level Indicators (SLIs) and SLOs, minimizing alert fatigue while maximizing system reliability.
  • Test Automation: Strong background in integrating automated test frameworks into CI/CD pipelines for continuous integration of AI features.
  • Prompt & Model Analysis: Analytical mindset to benchmark prompt variants, track regression in model behavior, and profile latency across model providers.

Programming & Operations

  • Python Mastery: Advanced Python programming skills, including experience with async execution, API integration, and AI frameworks (e.g., LangChain, LlamaIndex).
  • Production Operations: Solid understanding of cloud infrastructure (AWS/Google Cloud Platform/Azure), containerization (Docker, Kubernetes), and DevOps best practices.

 

Preferred Qualifications

  • 3+ years of experience operationalizing LLMs or generative AI applications in production.
  • Experience managing vector databases (e.g., Pinecone, Milvus, Chroma) and tracking RAG pipeline performance.
  • Background in Site Reliability Engineering (SRE) or specialized MLOps roles.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91131106
  • Position Id: 9103253
  • Posted 6 hours ago

Company Info

About Rivago infotech inc

Rivago Infotech Inc has been a leader in IT staffing and Software development for over 5 years and is one of the largest diversity and development firms in the industry. We are known for our high-touch, customer-eccentric approach, offering our clients unmatched quality, responsiveness and flexibility . We are appreciated by our clients for our streamlined execution, highly efficient service and exceptional talent management that go above and beyond traditional staffing services.

About_Company_OneAbout_Company_Two
Contact the job poster
RA

Rajat Arora

Recruiter @ Rivago infotech inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Charlotte, North Carolina

•

Today

Easy Apply

Contract, Third Party

60 - 63

Woodbridge Township, New Jersey

•

Today

Easy Apply

Full-time, Third Party

120,000 - 130,000

Paramus, New Jersey

•

Today

Easy Apply

Third Party, Contract

70 - 75

Hybrid in Dallas, Texas

•

Today

Easy Apply

Third Party, Contract

50 - 60

Search all similar jobs