AI Architect - Remote

Remote • Posted 3 hours ago • Updated 3 hours ago
Contract Corp To Corp
Contract W2
Contract Independent
12 Months
No Travel Required
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • AI Architect
  • Architect
  • A/B Testing
  • API
  • Amazon Web Services
  • Artificial Intelligence
  • Auditing
  • Authentication
  • Authorization
  • Build Vs Buy
  • Business Continuity Planning
  • Caching
  • Change Management
  • Cloud Architecture
  • Cloud Computing
  • Cloud Security
  • Communication
  • Computer Networking
  • Continuous Delivery
  • Continuous Integration
  • Data Centers
  • Database
  • Disaster Recovery
  • Documentation
  • Dynatrace
  • Enterprise Architecture
  • FOCUS
  • Failover
  • Good Clinical Practice
  • Google Cloud Platform
  • HIPAA
  • High Availability
  • Hosting
  • IDPS
  • IT Management
  • IaaS
  • Leadership
  • Load Balancing
  • Lua
  • Machine Learning (ML)
  • Management
  • Mentorship
  • Microservices
  • Microsoft
  • Microsoft Azure
  • Microsoft Certified Professional
  • Open Source
  • Orchestration
  • Privacy
  • Provisioning
  • Python
  • Reasoning
  • Regulatory Compliance
  • Release Management
  • Roadmaps
  • Routing
  • SaaS
  • Semantics
  • Servers
  • Solution Architecture

Summary

Role Summary

We are seeking an experienced AI Architect to design, govern, and scale enterprise-grade agentic AI solutions across hybrid and multi-cloud environments. The ideal candidate will bring deep expertise in architecting distributed AI systems that span on-premises infrastructure and multiple public clouds (AWS, Azure, Google Cloud Platform), with a strong focus on AI Gateway design, agent orchestration, model interoperability, and secure, compliant AI deployment at scale. This role sits at the intersection of enterprise architecture, cloud infrastructure, and applied AI translating business needs into resilient, secure, and cost-efficient AI platforms that support autonomous agents, LLM-powered applications, and multi-model workflows.

Key Responsibilities

Architecture & Design

  • Design end-to-end architecture for agentic AI systems (multi-agent orchestration, tool-use frameworks, memory/state management, planning and reasoning loops) deployed across hybrid and multi-cloud environments.
  • Architect and implement an AI Gateway layer to unify access to multiple LLM/model providers (OpenAI, Anthropic, Google, open-source models, self-hosted models) with centralized routing, rate limiting, load balancing, caching, and failover.
  • Define reference architectures for hybrid cloud AI deployments, balancing workloads across on-prem, private cloud, and public cloud (AWS/Azure/Google Cloud Platform) based on data residency, latency, cost, and compliance requirements.
  • Establish patterns for model orchestration, agent-to-agent communication, and tool/function calling across distributed systems.
  • Design for interoperability across cloud-native AI services (Bedrock, Azure AI Foundry, Microsoft Co-pilot)
  • Integrating MCP Registry, governance of Agent Registry
  • Microsoft EntraID as identity provider for MCP authentication
  • Integrating AI gateway with AWS Valkey cache and other vector / RAG databases
  • Specific implementation experience of Kong AI gateway with plug-in based architecture, ability create plug-ins using Lua script or Python
  • Observability integration with Dynatrace for agents and AI gateway telemetry; Splunk integration for audit logs

AI Gateway & Governance

  • Own the strategy and implementation of the AI Gateway as the control plane for all AI/LLM traffic including authentication, authorization, usage metering, cost governance, PII/data redaction, prompt/response logging, and audit trails.
  • Implement guardrails for model governance: version control, A/B testing, fallback routing, semantic caching, and observability across multiple model providers.
  • Define policies for responsible AI, data privacy, and security across the gateway and agentic workflows, ensuring compliance with regulatory frameworks (GDPR, SOC 2, HIPAA, etc. as applicable).

Multi-Cloud & Infrastructure

  • Architect resilient, scalable, and cost-optimized infrastructure spanning multiple cloud providers and on-prem data centers.
  • Drive infrastructure-as-code practices (Terraform, Pulumi) for consistent multi-cloud provisioning of AI/ML infrastructure.
  • Evaluate and integrate vector databases, feature stores, and RAG pipelines across hybrid environments.
  • Ensure high availability, disaster recovery, and business continuity for mission-critical agentic AI applications.

Leadership & Collaboration

  • Partner with engineering, product, security, and data teams to align AI architecture with business objectives.
  • Provide technical leadership and mentorship to AI/ML engineers and platform teams.
  • Evaluate emerging tools, frameworks, and vendors in the agentic AI and AI infrastructure ecosystem (e.g., MCP Registry, Kong AI Gateway, Agent365).
  • Create architecture documentation, design standards, and best practices to guide organization-wide AI adoption.
  • Act as a trusted advisor to leadership on AI platform strategy, build-vs-buy decisions, and technology roadmaps.
  • Setup Operating model to operate and govern AI gateway usage enterprise-wide
  • Scalability, hosting topology, define centralized vs federated model
  • Deployment architecture, change management, release management and incident management
  • Alignment with United AI governance Framework for agentic systems
  • Integrate skills, tools, MCP servers - internal or SaaS vendor source

Required Qualifications

  • 8+ years of experience in software/solutions architecture, with 3+ years focused on AI/ML systems.
  • Proven hands-on experience architecting agentic AI systems multi-agent frameworks, autonomous workflows, tool/function calling, and orchestration.
  • Strong expertise in hybrid and multi-cloud architecture (AWS, Azure, Google Cloud Platform, and on-premises/private cloud integration).
  • Demonstrated experience designing or implementing an AI Gateway / LLM Gateway(e.g., Kong, custom-built gateways) for managing multi-model access, routing, rate limiting, and observability.
  • Deep understanding of LLM ecosystems (OpenAI, Anthropic, open-source models) and model serving infrastructure.
  • Experience with RAG architectures, vector databases (e.g.Milvus, pgvector), and knowledge retrieval systems.
  • Solid grounding in API design, microservices, event-driven architecture, and distributed systems.
  • Strong knowledge of cloud security, IAM, networking, and compliance frameworks across multi-cloud environments (IDPs like EntraID and Ping.
  • Proficiency with infrastructure-as-code (Terraform/Harness, DeclarativeKong), CI/CD pipelines, and observability tooling (Splunk, DynaTrace).
  • Programming proficiency in Python (required).
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10339744
  • Position Id: 9071731
  • Posted 3 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

4d ago

Easy Apply

Contract, Third Party

Depends on Experience

Remote

7d ago

Easy Apply

Contract

Depends on Experience

Remote or Irving, Texas

30+d ago

Easy Apply

Contract

70 - 75

Remote or Lombard, Illinois

Today

Full-time

USD 125,000.00 per year

Search all similar jobs