Lead the design and engineering evolution of enterprise AI platform capabilities including AI gateways, model access and serving, model routing, RAG, AI agents, tool execution, orchestration, evaluation, observability, and LLMOps/MLOps.
Solve complex engineering trade-offs involving latency, throughput, resiliency, scalability, security, data isolation, portability, and cost.
Lead technical spikes, prototypes, reference implementations, deep design reviews, performance analysis, and production troubleshooting.
Establish engineering standards for availability, recovery, performance, capacity, telemetry, release safety, evaluation coverage, and inference cost.
Build reusable platform assets such as APIs, SDKs, Terraform/IaC modules, deployment patterns, CI/CD templates, dashboards, evaluation frameworks, and developer tooling.
Drive production excellence through observability, traceability, controlled releases, rollback strategies, incident learning, capacity planning, and cost optimization.
Implement AI security controls including IAM, authorization-aware retrieval, secure tool execution, prompt-injection defenses, data protection, logging, and auditability.
Partner with Cybersecurity, Risk, Compliance, Legal, Audit, Architecture, Product, and Business teams to translate enterprise requirements into practical technical controls.
Evaluate emerging AI technologies through hands-on technical assessments and determine adoption based on value, maturity, risk, operational fit, and total cost of ownership.
Mentor senior engineers and drive adoption of enterprise AI platform patterns across multiple engineering teams.
10+ years of progressive experience in software engineering, distributed systems, cloud/platform engineering, AI/ML infrastructure, or related technical domains.
Proven experience building, scaling, transforming, or troubleshooting production platforms used by multiple engineering teams, products, or business domains.
Strong hands-on experience in several of the following:
LLM / Generative AI platforms
Model serving and inference
Inference optimization
AI gateways
Model routing
RAG
Embeddings / Vector Search
AI Agent frameworks
Tool execution / MCP
AI orchestration
LLMOps / MLOps
AI evaluation frameworks
AI observability
AI security / guardrails
Strong foundation in distributed systems, API/platform design, cloud-native architecture, containers, Kubernetes, networking, IAM, and secrets management.
Experience developing reusable engineering patterns, frameworks, APIs, infrastructure modules, CI/CD pipelines, or developer platforms.
Demonstrated ability to provide technical leadership across multiple teams without direct reporting responsibility.