Job Summary The Machine Learning Engineer Lead will define and lead the architecture of scalable AI/ML and agentic systems across the global product portfolio. This senior technical leadership role will focus on large-scale distributed ML systems, LLM and RAG architectures, Agentic AI frameworks, tool orchestration, and enterprise platform engineering. The role will shape long-term AI platform strategy, establish technical standards, and mentor engineering teams while supporting highly available, secure, and scalable AI systems. Key Responsibilities Define reference architectures for LLM, machine learning, and agent-based systems across products. Design high-availability and low-latency inference platforms capable of operating at global scale. Establish reusable platform components for model lifecycle management, deployment, monitoring, and operationalization. Architect scalable AI platforms supporting large-scale distributed ML systems and enterprise AI workloads. Architect multi-step, reasoning-driven agentic AI systems. Design orchestration patterns for tool use, API invocation, and structured function calling. Lead the implementation and governance of Model Context Protocol (MCP) servers to standardize tool integration and context management. Define guardrails, permissions, security controls, and audit mechanisms for enterprise-safe AI systems. Establish and maintain best practices for MLOps, CI/CD, observability, scalability, and system reliability. Design and implement scalable inference systems using containerization and Kubernetes. Drive the deployment and optimization of LLM, Generative AI, and RAG solutions in production environments. Design cloud-based AI/ML architectures across AWS, Azure, or Google Cloud Platform. Establish technical standards and architectural patterns for AI/ML and agentic systems across engineering teams. Embed Responsible AI principles into platform architecture and engineering practices. Provide technical leadership, mentorship, and guidance to senior engineers and engineering teams. Collaborate with cross-functional teams to influence technical direction and ensure alignment with enterprise AI platform strategy. Support people management, leadership, and team development activities as required. Required Qualifications 10+ years of experience building and deploying production-grade machine learning systems at scale. Strong experience with LLMs, Generative AI, and RAG deployments in production environments. Strong Python development background. Expertise designing and implementing AI/ML systems in cloud environments such as AWS, Azure, or Google Cloud Platform. Hands-on experience with Kubernetes, containerization, and scalable inference systems. Experience designing agentic AI systems and tool orchestration frameworks. Experience implementing and governing MCP servers or structured architectures for tool integration and context management. Experience with large-scale distributed ML systems and enterprise platform engineering. Experience establishing MLOps, CI/CD, observability, and system reliability practices. Demonstrated people management, technical leadership, or mentorship experience. Strong understanding of high-availability and low-latency AI/ML architectures. Ability to define technical standards and influence architecture and engineering decisions across teams. Education: Bachelors Degree
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: compun
- Position Id: ALIDC5872695
- Posted 4 hours ago