Role: Real-Time Inference Engineering Lead
Location: Charlotte, NC — Onsite
Employment Type: Full-Time
About the Role
We are seeking a Real-Time Inference Engineering Lead to design, build, and industrialize low-latency, resilient model-serving services for real-time predictive AI use cases.
The Cortex Predictive AI Platform accelerates predictive AI modernization and enterprise adoption across the full model lifecycle, including governed data and features, model development and validation, deployment and inference, and ongoing monitoring and operations.
As part of the Real-Time Services portfolio, this role will provide technical leadership for online inference architecture, deployment patterns, API services, capacity controls, observability, reliability, and operational practices across public cloud and on-premises environments.
This is a hands-on engineering leadership role requiring strong expertise in model serving, Kubernetes, APIs, distributed systems, performance optimization, reliability engineering, CI/CD, and production operations.
Key Responsibilities
- Define target architecture, engineering standards, and reusable patterns for real-time predictive model-serving services across cloud and on-premises environments.
- Design, build, test, deploy, and operate scalable online inference services that meet stringent latency, throughput, availability, resiliency, security, and reliability requirements.
- Establish reusable serving patterns for:
- Synchronous inference APIs
- Asynchronous inference
- Batch-adjacent processing
- Event-driven real-time use cases
- Develop standardized model deployment approaches covering model packaging, versioning, release promotion, canary deployment, rollback, and retirement.
- Design and implement secure inference APIs, including authentication, authorization, traffic management, rate limiting, auditability, and enterprise integration.
- Engineer Kubernetes-based serving platforms using GKE, OpenShift, and related container orchestration technologies.
- Implement autoscaling, resource allocation, quota management, capacity planning, workload isolation, and other controls for variable inference demand.
- Lead performance engineering, load testing, stress testing, benchmarking, profiling, and failure testing for production inference workloads.
- Identify and implement latency-optimization opportunities across model initialization, feature retrieval, networking, API processing, runtime configuration, and infrastructure utilization.
- Define and implement monitoring and telemetry standards, including dashboards, alerts, distributed tracing, SLIs, SLOs, and error-budget practices.
- Partner with ML platform, data engineering, application engineering, security, networking, and operations teams to integrate model services with governed data, feature, identity, and infrastructure capabilities.
- Implement CI/CD pipelines and automated validation for model-serving services, infrastructure configuration, APIs, performance benchmarks, and release-readiness controls.
- Develop operational runbooks, incident-response procedures, support models, disaster-recovery plans, and production-readiness artifacts.
- Drive reliability improvements through root-cause analysis, capacity reviews, resiliency testing, disaster-recovery exercises, and continuous operational improvement.
- Mentor engineers and establish reusable technical documentation, reference implementations, engineering standards, and knowledge-transfer materials.
Required Qualifications
- 8+ years of experience in software engineering, platform engineering, cloud engineering, SRE, infrastructure engineering, or a related discipline.
- 4+ years designing, building, or operating production APIs, distributed systems, platform services, or real-time data/ML workloads.
- Demonstrated experience providing technical leadership for the architecture and delivery of highly available, performance-sensitive production services.
- Strong hands-on experience with online inference architecture, model-serving frameworks, or predictive-model deployment patterns.
- Strong experience designing and operating RESTful, gRPC, or event-driven APIs.
- 4+ years of strong production experience with Kubernetes and container platforms, including GKE, OpenShift, or comparable environments.
- Experience with autoscaling, resource management, capacity planning, performance testing, load testing, and optimization of distributed services.
- Experience implementing observability practices, including monitoring, dashboards, alerts, distributed tracing, SLIs, SLOs, and incident management.
- Strong experience with CI/CD, Git-based development, automated testing, deployment automation, and production-release processes.
- Strong understanding of resiliency, high availability, fault tolerance, disaster recovery, and operational support for critical production services.
- Ability to collaborate effectively with data science, ML engineering, platform engineering, application engineering, security, operations, and business stakeholders.
Required Technical Skills
Model Serving & Online Inference
- Online inference and low-latency model-serving architecture
- Model deployment, packaging, versioning, routing, rollout, rollback, and lifecycle management
- Predictive-model productionization and serving patterns
- Model performance optimization
APIs & Distributed Systems
- REST APIs and gRPC
- API gateways
- Authentication and authorization
- Traffic management and rate limiting
- Event-driven architectures
- Distributed systems and service-to-service communication
- API observability and operational controls
Kubernetes & Cloud Platforms
- Kubernetes and containers
- GKE and/or OpenShift
- Ingress and service networking
- Service mesh
- Workload scheduling and isolation
- Horizontal/vertical autoscaling
- Resource quotas and capacity controls
Performance & Reliability Engineering
- Load and stress testing
- Benchmarking and profiling
- Latency optimization
- Capacity planning
- Fault injection and resiliency testing
- High availability and fault tolerance
- Disaster recovery
- Root-cause analysis and incident response
Observability
- Metrics, logs, and distributed tracing
- Monitoring and telemetry
- Dashboards and alerting
- SLIs and SLOs
- Error budgets
- Production health and readiness monitoring
DevOps & Automation
- CI/CD
- Git-based development workflows
- Automated testing
- Infrastructure as Code
- Deployment automation
- Release controls and promotion strategies
Preferred Qualifications
- Experience with Vertex AI Endpoints, KServe, Seldon, NVIDIA Triton Inference Server, MLflow deployments, or comparable model-serving technologies.
- Experience deploying and operating models on Google Cloud Platform, AWS, Azure, private cloud, or hybrid-cloud environments.
- Experience with service mesh, API gateway, traffic-routing, edge-serving, or distributed proxy technologies.
- Experience serving high-volume, customer-facing, fraud, risk, personalization, decisioning, or other latency-sensitive predictive models.
- Experience with feature serving, online feature stores, caching, streaming platforms, or real-time data enrichment.
- Experience with Terraform, Helm, Argo CD, Jenkins, GitHub Actions, GitLab CI, or similar automation technologies.
- Experience working in banking, financial services, healthcare, insurance, or another regulated enterprise environment.
- Experience participating in a 24x7 operational support model for high-priority production services.
Expected Outcomes
The successful candidate will help establish and deliver:
- Standardized, production-ready real-time inference architecture and reusable model-serving patterns.
- Reliable online inference services that consistently meet defined latency, throughput, availability, and resiliency objectives.
- Automated model deployment, testing, monitoring, capacity management, and rollback capabilities.
- Clear operational dashboards, SLOs, alerts, runbooks, and production-readiness evidence.
- Improved engineering productivity and accelerated adoption of secure, scalable real-time predictive AI capabilities across the Cortex portfolio.