Job Role: Real - Time Inference Engineering Lead.
Location: Charlotte, NC
Duration: Long-Term Contract
Job Summary
We are seeking a Real-Time Inference Engineering Lead to design, build, and operationalize scalable, low-latency model-serving platforms supporting real-time AI and predictive decisioning use cases.
The ideal candidate will have strong experience in AI/ML platform engineering, real-time inference, distributed systems, Kubernetes, MLOps, cloud/on-premises infrastructure, and production reliability. This role will lead architecture, deployment, performance optimization, resiliency, and operational governance for enterprise-scale online inference services.
Key Responsibilities
- Design and implement highly available, low-latency model-serving architectures for real-time inference workloads.
- Develop standardized deployment patterns for scalable AI/ML services using Kubernetes and containerized platforms.
- Lead the design, integration, and lifecycle management of API-based inference services.
- Optimize model-serving performance through latency tuning, caching, autoscaling, and resource management.
- Establish monitoring, observability, logging, tracing, SLOs/SLAs, alerting, and production runbooks.
- Lead load testing, capacity planning, performance engineering, resiliency, and disaster recovery initiatives.
- Integrate inference platforms with CI/CD pipelines for automated and controlled deployments.
- Partner with Data Science, MLOps, Platform Engineering, DevOps, and Infrastructure teams to ensure reliable production model operations.
- Define and govern enterprise standards for security, reliability, performance, and operational excellence across AI deployments.
- Provide technical leadership and mentorship for real-time AI platform engineering initiatives.
Required Skills
- Strong experience with online inference and real-time model-serving architectures.
- Strong knowledge of real-time APIs and distributed systems design.
- Hands-on experience with Kubernetes, containers, orchestration, and service mesh technologies.
- Experience with autoscaling, capacity management, workload optimization, and resource management.
- Strong background in performance engineering, load testing, and latency optimization.
- Experience with monitoring, observability, logging, tracing, and SLO management.
- Strong understanding of CI/CD, DevOps, and Infrastructure-as-Code (IaC) practices.
- Experience designing reliable, fault-tolerant, and resilient production systems.
- Experience supporting cloud and on-premises platforms.
- Strong programming experience with Python, Java, Go, or similar backend technologies.
- Experience with MLOps platforms and enterprise AI/ML deployment frameworks.
Preferred Qualifications
- Hands-on experience with real-time recommendation, fraud detection, risk, personalization, or predictive analytics platforms.
- Experience with GPU-based inference and model optimization.
- Experience with multi-cloud AI/ML deployments.
- Experience operating AI/ML platforms at enterprise scale.
- Strong leadership experience working across Data Science, MLOps, Infrastructure, Platform Engineering, and DevOps teams.
Ideal Candidate
The successful candidate will combine AI platform engineering, distributed systems expertise, Kubernetes, MLOps, and operational excellence to deliver resilient, scalable, and high-performance real-time inference platforms for enterprise AI solutions.
Regards
Mark R
Email: