Enterprise AI Platform Design Lead - GenAI / Agentic AI
ROLE OVERVIEW
We are seeking a Senior AI Platform Design Architect to lead the architecture and design of a secure, enterprise-scale on-premises AI platform for a major financial institution. The architect will define the target-state architecture for running Generative AI, LLM and Agentic AI workloads within an enterprise-controlled environment, spanning GPU infrastructure, AI/ML platforms, model serving, data, networking, security, governance, observability and enterprise integration.
KEY RESPONSIBILITIES
Define the target architecture and technical blueprint for an on-premises Enterprise AI Platform.
Design infrastructure supporting LLM inference/model serving, Generative AI applications, Agentic AI, RAG and AI/ML workloads.
Define GPU compute architecture, capacity planning, workload allocation and scalability strategies.
Design containerized AI platforms using Kubernetes, OpenShift or equivalent technologies.
Define model-serving patterns for hosting enterprise-approved foundation models.
Develop reusable AI platform reference architectures, standards and design patterns.
Design AI/LLM gateway architecture for model routing, access control, policy enforcement, usage monitoring and cost management.
Design Agent/Tool integration patterns, including APIs and MCP/A2A where applicable.
Define enterprise RAG architecture covering ingestion, embeddings, vector databases, retrieval and knowledge governance.
Establish AI security, identity, access control, data protection and network-segmentation architecture.
Define observability, logging, monitoring, evaluation and operational telemetry for AI workloads.
Establish DevSecOps, MLOps and LLMOps architecture and deployment patterns.
Define high availability, disaster recovery, backup and business-continuity architecture.
Partner with Enterprise Architecture, Cybersecurity, Infrastructure, Data and AI/ML teams to establish platform standards.
Create architecture diagrams, technical specifications, ADRs and implementation roadmaps; conduct architecture reviews.
CORE TECHNICAL EXPERIENCE
DOMAIN REQUIRED / PREFERRED EXPERIENCE
On-Prem Infrastructure Large-scale enterprise on-prem platforms; data center compute/storage/networking; GPU infrastructure; VMware/OpenStack or equivalent; HA/DR.
AI / GenAI Generative AI, LLM inference/model serving, Agentic AI, RAG, vector databases, embeddings, MLOps/LLMOps, evaluation and guardrails.
Containers & Platform Kubernetes/OpenShift, Docker, Helm/operators, service mesh, API gateways, CI/CD, DevSecOps, Terraform or equivalent IaC.
Enterprise Integration REST APIs, microservices, event-driven architecture, enterprise application integration, databases and data platforms.
Security IAM/RBAC, Zero Trust, network segmentation, encryption, secrets management, data protection and AI security/governance.
Cloud - Secondary Azure/AWS/Google Cloud Platform experience, especially hybrid architecture and adapting cloud-native AI patterns to on-prem environments.
PREFERRED BACKGROUND
Enterprise AI platform / AI Landing Zone / Private AI architecture.
On-prem LLM infrastructure and NVIDIA GPU ecosystem.
Kubernetes/OpenShift AI platforms; Red Hat OpenShift AI, NVIDIA AI Enterprise or comparable technologies.
Model-serving technologies such as vLLM, NVIDIA Triton or equivalent.
Enterprise RAG and knowledge platforms.
Large-scale financial-services, banking or other highly regulated environments.
KEY ARCHITECTURE DELIVERABLES
On-Prem Enterprise AI Platform Reference Architecture
GPU & Compute Architecture
Kubernetes / Container Platform Architecture
LLM Model Hosting & Inference Architecture
Enterprise RAG Architecture
AI / Agent Integration Architecture
AI Security & Governance Architecture
Data, Storage & Network Architecture
AI Observability & Operations Architecture
HA/DR & Resilience Architecture
DevSecOps / MLOps / LLMOps Architecture
Platform Capacity, Scalability Model & Implementation Roadmap
IDEAL CANDIDATE PROFILE
The ideal candidate combines Enterprise Architecture, AI Architecture and Infrastructure/Platform Engineering. They should be able to design the complete stack from AI applications and agents through the AI/LLM platform, model serving, GPU/Kubernetes, data, storage, networking, security and operations and translate the architecture into an implementable production roadmap.
AI APPLICATIONS AGENTS AI/LLM PLATFORM MODEL SERVING GPU/KUBERNETES DATA SECURITY OPERATIONS
New York-based candidates strongly preferred.