Real-Time Inference Engineering Lead

Concord, CA, US • Posted 4 days ago • Updated 1 hour ago
Contract Corp To Corp
Contract W2
Contract Independent
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • REAL-TIME
  • AI
  • MLops
  • Inference Engineering

Summary

Real-Time Inference Engineering Lead

The Real-Time Inference Engineering Lead is responsible for designing, building, and operationalizing scalable, low-latency model serving platforms that power real-time AI and predictive decisioning use cases. This role leads the architecture, deployment, performance optimization, resiliency, and operational governance of online inference services across cloud and on-premises environments.

* Design and implement highly available, low-latency model serving architectures for real-time inference workloads.
* Develop and standardize deployment patterns for scalable AI/ML services across Kubernetes-based platforms.
* Lead API-based inference service design, integration, and lifecycle management.
* Optimize model serving performance through latency tuning, caching strategies, autoscaling, and resource management.
* Establish monitoring, observability, SLOs/SLAs, alerting, and operational runbooks for production services.
* Drive load testing, capacity planning, resiliency engineering, and disaster recovery readiness.
* Integrate inference platforms with CI/CD pipelines to enable automated deployments and controlled releases.
* Partner with Data Science, Platform Engineering, MLOps, and Infrastructure teams to ensure reliable production model operations.
* Govern operational best practices, security, reliability, and performance standards for enterprise AI deployments.

* Online inference and model-serving architectures
* Real-time APIs and distributed systems design
* Kubernetes, container orchestration, and service mesh technologies
* Autoscaling, capacity management, and workload optimization
* Performance engineering, load testing, and latency optimization
* Monitoring, observability, logging, tracing, and SLO management
* CI/CD, DevOps, and Infrastructure-as-Code practices
* Reliability engineering, fault tolerance, and resiliency patterns
* Cloud and on-premises platform operations
* Python, Java, Go, or similar backend development experience

* Experience with MLOps platforms and enterprise AI deployment frameworks.
* Hands-on experience with real-time recommendation, fraud, risk, personalization, or predictive analytics platforms.
* Familiarity with GPU-based inference, model optimization, and multi-cloud deployments.

A successful candidate combines AI platform engineering, distributed systems expertise, and operational excellence to deliver resilient, high-performance inference platforms that enable enterprise-scale real-time AI solutions.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91009966
  • Position Id: 2026-30889
  • Posted 4 days ago
Contact the job poster
Maheshwaran Gurusamy

Maheshwaran Gurusamy

Associate Client Partner @ Cloudious
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

San Francisco, California

Today

Full-time

USD 160,000.00 - 250,000.00 per year

San Francisco, California

Today

Full-time

USD 200,000.00 - 290,000.00 per year

Santa Clara, California

10d ago

Easy Apply

Third Party, Contract

Depends on Experience

Palo Alto, California

Today

Full-time

USD 135,000.00 - 175,000.00 per year

Search all similar jobs