Inference Lead

Charlotte, NC, US • Posted 1 day ago • Updated 1 hour ago
Contract W2
Contract Independent
Contract Corp To Corp
6 Months
On-site
$DOE
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Artificial Intelligence
  • Use Cases
  • Load Balancing
  • Benchmarking
  • Performance Tuning
  • Incident Management
  • Failover
  • Disaster Recovery
  • Mentorship
  • Real-time
  • Microservices
  • Management
  • Performance Engineering
  • Optimization
  • Load Testing
  • Capacity Management
  • Continuous Integration
  • Continuous Delivery
  • High Availability
  • Computer Science
  • Cloud Computing
  • Kubernetes
  • Machine Learning (ML)
  • IMG
  • LinkedIn
  • SAINT
  • Technical Direction

Summary

Job Title:- Inference Lead (Machine Learning Platform Engineer Lead Real-Time Inference)

Location:- Charlotte, North Carolina (Hybrid Onsite - local candidates preferred. )

Duration:- 6 months

All visa except OPT/CPT

Contract

Real-Time Services Real-Time Inference Engineering Lead

Role Summary

The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments.

Key Responsibilities

  • Architect low-latency online inference and real-time model-serving solutions.
  • Develop scalable APIs, microservices, and deployment patterns for predictive models.
  • Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
  • Conduct benchmarking, performance tuning, capacity planning, and load testing.
  • Optimize latency, throughput, resource consumption, availability, and cost.
  • Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
  • Build CI/CD pipelines for repeatable model and service releases.
  • Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
  • Lead technical reviews and mentor inference and platform engineers.

Required Skills

  • Online inference and real-time model-serving architecture.
  • REST/gRPC APIs and distributed microservices.
  • Kubernetes, containers, autoscaling, and traffic management.
  • Performance engineering, latency optimization, and load testing.
  • Monitoring, SLOs, capacity planning, and production operations.
  • CI/CD and progressive-deployment approaches.
  • Resilience and high-availability engineering.
  • Cloud and on-premises deployment experience.

Preferred Qualifications

  • Degree in computer science, engineering, or a related discipline.
  • Experience with enterprise model-serving platforms and inference runtimes.
  • Cloud, Kubernetes, SRE, or ML engineering certification.

Skill

Years of experience

Last used/Worked (year)

Candidate self-rating(out of 10)

Navya Gupta
Sr. IT Technical Recruiter

Email:

Gtalk:
Phone: +1

Linkedin id:
Address: 505 Knolle Court, Saint Augustine| FL 32092

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91022079
  • Position Id: 2026-51087
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Charlotte, North Carolina

Yesterday

Easy Apply

Contract

$88.85

Charlotte, North Carolina

Today

Full-time

USD 70.00 - 80.00 per hour

Charlotte, North Carolina

Today

Full-time

USD 83,520.00 - 125,280.00 per year

Hybrid in Charlotte, North Carolina

8d ago

Easy Apply

Contract, Third Party

$80

Search all similar jobs