Hello Everyone,
Position: #MLops_Engineer/ #AI_Inference_Engineering_Lead
Location: Charlotte, NC
Note: Mention Visa and Current location when you reach out to me.
Description:
The Real-Time Inference Engineering Lead is responsible for designing, building, and operationalizing scalable, low-latency model serving platforms that power real-time AI and predictive decisioning use cases. This role leads the architecture, deployment, performance optimization, resiliency, and operational governance of online inference services across cloud and on-premises environments.
· Design and implement highly available, low-latency model serving architectures for real-time inference workloads.
· Develop and standardize deployment patterns for scalable AI/ML services across Kubernetes-based platforms.
· Lead API-based inference service design, integration, and lifecycle management.
· Optimize model serving performance through latency tuning, caching strategies, autoscaling, and resource management.
· Establish monitoring, observability, SLOs/SLAs, alerting, and operational runbooks for production services.
· Drive load testing, capacity planning, resiliency engineering, and disaster recovery readiness.
· Integrate inference platforms with CI/CD pipelines to enable automated deployments and controlled releases.
· Partner with Data Science, Platform Engineering, MLOps, and Infrastructure teams to ensure reliable production model operations.
· Govern operational best practices, security, reliability, and performance standards for enterprise AI deployments.
· Online inference and model-serving architectures
· Real-time APIs and distributed systems design
· Kubernetes, container orchestration, and service mesh technologies
· Autoscaling, capacity management, and workload optimization
· Performance engineering, load testing, and latency optimization
· Monitoring, observability, logging, tracing, and SLO management
· CI/CD, DevOps, and Infrastructure-as-Code practices
· Reliability engineering, fault tolerance, and resiliency patterns
· Cloud and on-premises platform operations
· Python, Java, Go, or similar backend development experience
· Experience with MLOps platforms and enterprise AI deployment frameworks.
· Hands-on experience with real-time recommendation, fraud, risk, personalization, or predictive analytics platforms.
· Familiarity with GPU-based inference, model optimization, and multi-cloud deployments.
Regards