Full time, remote opportunity for strong Senior Architect, Google Cloud and AI Platform Architecture. Permanent residents only. No agenices please.
The Senior Architect, Google Cloud Platform & AI Platform Architect is a senior individual-contributor architecture role. This person will architect and evolve platform capabilities for model serving, infrastructure fit, Model Gateway and tool integration, MLOps, observability, cost-per-unit economics, and peak-load readiness. It collaborates with cross-functional stakeholders and peer senior architects in Design and Stability, and sets platform architecture direction so systems are scalable, performant, reliable, and cost-efficient. At least fifty percent of this role is hands-on architecture work, including platform design, infrastructure decisions, MLOps patterns, reviews, and authoring standards.
RESPONSIBILITIES:
Strategic Platform Leadership
§ Translate enterprise AI strategy defined by the AI Architect Principal, Enterprise Architecture, Enterprise AI Tech Leaders (AI Engineering, DS/ML Engineering, AI Emerging Tech), AI Product Teams into actionable platform roadmaps and technical priorities.
§ Partner with data, cloud, security, and infrastructure teams to define the end-to-end AI architecture framework, including compute, model lifecycle, and deployment strategies.
§ Evaluate emerging technologies and recommend platform enhancements to improve model performance, scalability, and sustainability.
§ Define and deliver AI/ML/Agentic operations strategy – including tool suite, standards, and technology/capability roadmap.
§ Establish POV on key evaluations – including, but not limited to “buy v. build”, platform assessments, tool comparisons, cost/performance optimizations
§ Drive the technical strategy for the platform, balancing short-term needs with long-term scalability and reliability.
§ Design, implement, and scale cloud-based infrastructures (Google Cloud Platform and Databricks) to support internal and external applications.
§ Oversee platform architecture decisions, ensuring that the platform is robust, efficient, and capable of supporting growing business needs.
Architecture Design and Governance
§ Lead the design of core AIML platform components—data pipelines, model training and inference engines, orchestration workflows, and monitoring frameworks.
§ Design for continuous training pipelines and promote automation capabilities where possible
§ Establish architectural best practices, patterns, and standards for AI/ML development and deployment.
§ Oversee design reviews and ensure compliance with enterprise architecture and regulatory requirements.
§ Design and build platform to meet AI Governance requirements (including application of controls and measurement of controls). Ensure observability requirements can be met.
§ Serve as a standing member of the Architecture Review Board for the Scale lane, advising on infrastructure fit, serving readiness, and cost-per-unit implications.
Cross-Functional Collaboration
§ Work closely with product, data science, engineering, and security teams to operationalize AI/ML capabilities across the enterprise.
§ Partner with the AI Architect Principal and AI Solution Architects to align agentic platform evolution with organizational priorities and technology roadmaps.
§ Collaborate with security and operations teams to ensure the platform is secure, compliant, and maintains high uptime and reliability.
§ Engage with leadership to align technical initiatives with organizational objectives.
Architecture Mentorship and Technical Enablement
§ Mentor AI Architects, AI engineers, and AI Ops engineers through design reviews, platform guidance, and technical coaching, raising platform quality through influence and reusable patterns
§ Foster a high-performance culture emphasizing innovation, collaboration, and agile delivery.
§ Provide hands-on technical leadership, leading by example in designing and building robust, scalable systems.
§ Drive high standards of code quality, testing, and engineering practices.
§ Advocate for platform engineering best practices, ensuring systems are maintainable, extensible, and documented.
Team Leadership and Talent Development
§ Manage and mentor a team of AI engineers, AI architects, and AI Ops engineers.
§ Foster a high-performance culture emphasizing innovation, collaboration, and agile delivery.
§ Support career development, technical upskilling, and diversity in AI technology roles, specifically those aligned into the Architecture space. Mentor and guide junior team members.
§ Conduct regular performance reviews, provide career development guidance, and support team members’ growth and skill development.
§ Own hiring, onboarding, and team-building activities to ensure the team has the right talent and skills.
§ Provide hands-on technical leadership, leading by example in designing and building robust, scalable systems.
§ Drive high standards of code quality, testing, and engineering practices.
§ Advocate for platform engineering best practices, ensuring systems are maintainable, extensible, and documented.
Operational Excellence
§ Oversee platform scalability, reliability, and cost optimization across cloud and on-prem environments.
§ Implement observability and monitoring tools to proactively identify performance or security issues.
§ Ensure platform compliance with responsible AI, data privacy, and ethical ML principles.
§ Build and develop architecture audit process, and execute
§ Identify opportunities for automation, process improvements, and tooling that enhance platform reliability and efficiency.
§ Stay current with emerging technologies and industry best practices and incorporate relevant trends into the platform strategy.
§ Lead incident response and post-mortem reviews to continuously improve platform resilience.
Platform Performance & Reliability:
§ Define, implement, and monitor key platform performance metrics, including system uptime, latency, and resource utilization.
§ Ensure the platform is scalable and cost-efficient, optimizing for performance and operational cost efficiency.
§ Lead efforts to identify, troubleshoot, and resolve platform performance issues or outages.
REQUIREMENTS:
§ Bachelor’s degree in computer science, a related field, or applicable work experience.
§ 10+ years of experience in software development or architecture; minimum of 3+ years of individual-contributor technical leadership in platform or architecture roles.
§ Deep knowledge of AI/ML ecosystems.
§ Experience designing MLOps pipelines and AI Ops frameworks at enterprise scale.
§ Strong expertise in Google Cloud Platform and cloud-native architecture, with hands-on knowledge of Google Cloud Platform services and architecture patterns supporting AI/ML platforms, MLOps frameworks, and CI/CD practices.
§ Experience with multi-agent architecture concepts: orchestration patterns, tool/skill registry, memory and state management, and agent observability
§ Experience setting technical standards and coordinating across distributed engineering teams
§ Proven ability to lead platform architecture across cross-functional engineering partners and deliver enterprise-scale AI platform solutions.
§ Comfortable navigating new, novel technology solutions and working through ambiguity to deliver in evolving domains
§ Familiarity with security and compliance standards for platform operations.
§ Strong analytical and problem-solving skills with a data-driven mindset.
§ Excellent communication, program coordination, and stakeholder engagement skills.
§ Comfortable with presenting up to senior leadership (VP+ level); able to present technical concepts to non-technical and executive levels.