Position Title: Principal AI Platform Engineer
Location: New York NY (Hybrid 3 days onsite)
Duration: 12+ months
Job Description:
The Client is seeking a Principal AI Platform Engineer / Agentic AI Architect to lead the architecture, development and production deployment of an enterprise-grade AI platform supporting financial services, insurance, payments and healthcare operations.
This position requires a deeply technical architect who can design the complete AI platform layer, remain hands-on with engineering and guide multiple product teams. The selected consultant will own LLM infrastructure, agent orchestration, retrieval systems, model evaluation, enterprise integrations, observability, security and Responsible AI controls.
The ideal candidate will have successfully taken multiple Generative AI and Agentic AI solutions from proof of concept into highly regulated production environments.
Key Responsibilities:
Define the end-to-end architecture for a multi-tenant enterprise Agentic AI platform.
Design and implement multi-agent systems containing planner, supervisor, retriever, executor, validator and human-in-the-loop components.
Build reusable agent frameworks using LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI or comparable orchestration technologies.
Architect a model gateway supporting Azure OpenAI, AWS Bedrock, Anthropic Claude, Google Gemini, Meta Llama and internally hosted models.
Design Retrieval-Augmented Generation and GraphRAG solutions across structured and unstructured enterprise data.
Implement hybrid retrieval combining vector search, semantic search, keyword search, reranking and knowledge graphs.
Build enterprise memory layers supporting episodic, semantic and procedural agent memory.
Design MCP-native tools, connectors and secure integration patterns for enterprise applications.
Integrate AI services with SAP, Oracle Fusion, Salesforce, ServiceNow, payment platforms and internal data products.
Design and operate LLM inference and model-serving infrastructure across cloud and on-premises environments.
Establish LLMOps and MLOps pipelines covering model registration, deployment, evaluation, monitoring, rollback and continuous improvement.
Develop automated evaluation frameworks for accuracy, groundedness, hallucination, safety, latency and cost.
Lead fine-tuning, instruction tuning, RLHF, RLAIF, model distillation and post-training initiatives.
Implement AI guardrails, content filtering, PII masking, access controls and prompt-injection defenses.
Establish model-risk documentation, lineage, explainability and audit evidence for regulated environments.
Design Kubernetes-based deployment architectures using Helm, Argo CD, Terraform and GitOps.
Implement distributed tracing and AI observability using OpenTelemetry, Grafana, Prometheus and specialized LLM monitoring platforms.
Define platform standards, reusable reference architectures and engineering best practices.
Lead architecture reviews and provide technical direction to AI engineers, platform engineers and product teams.
Partner with security, legal, risk, data governance and executive stakeholders.
Own technical decisions concerning build-versus-buy, model selection, infrastructure and platform scalability.
Optimize GPU utilization, token consumption, inference latency and overall LLM operating costs.
Maintain hands-on involvement through prototyping, code reviews, troubleshooting and production support.
Required Qualifications:
15+ years of overall software engineering, platform engineering or enterprise architecture experience.
8+ years designing and operating cloud-native platforms in production.
5+ years of hands-on AI/ML engineering or machine-learning platform experience.
3+ years building production Generative AI solutions using commercial or open-source LLMs.
Demonstrated experience architecting and deploying at least two enterprise Agentic AI implementations.
Expert-level Python development experience, including FastAPI, Flask, REST, gRPC and event-driven microservices.
Advanced experience with LangGraph and at least two additional agent-orchestration frameworks.
Hands-on experience with Azure OpenAI, AWS Bedrock and at least one additional foundation-model platform.
Deep knowledge of RAG, GraphRAG, embeddings, semantic search, reranking and context engineering.
Production experience with at least three vector technologies, including pgvector, Pinecone, Weaviate, Milvus, OpenSearch or Azure AI Search.
Advanced experience with Neo4j or another enterprise knowledge-graph platform.
Hands-on experience designing and implementing Model Context Protocol servers and clients.
Experience with LangSmith, MLflow and at least one evaluation framework such as Ragas, TruLens, Phoenix or Arize.
Strong Kubernetes engineering experience, including Helm, service meshes, autoscaling and production troubleshooting.
Advanced Infrastructure-as-Code experience using Terraform and strong GitOps experience using Argo CD.
Experience implementing model gateways, LLM routing, fallback strategies and multi-model architectures.
Hands-on experience deploying self-hosted LLMs using vLLM, NVIDIA Triton, Ray Serve or Hugging Face TGI.
Experience with GPU infrastructure, inference optimization, batching, quantization and distributed model serving.
Strong working knowledge of PostgreSQL, ClickHouse, object storage and distributed caching technologies.
Experience implementing OpenTelemetry-based tracing across agents, models, tools and microservices.
Strong understanding of OAuth 2.0, OIDC, workload identity, secrets management, encryption and zero-trust security.
Experience delivering AI platforms within banking, insurance, payments or another highly regulated industry.
Strong knowledge of Responsible AI, model governance, data privacy and model-risk management.
Experience presenting complex architectural decisions to executive and nontechnical stakeholders.
Proven ability to lead globally distributed engineering teams while remaining hands-on.
Mandatory Domain Experience:
Recent production experience in at least two of the following: banking and payment operations; insurance claims; fraud detection or financial-crime operations; healthcare operations; financial regulatory compliance; enterprise case management and workflow automation.
Highly Preferred Qualifications:
Experience building agentic workflows for claims adjudication, payment exceptions, fraud investigation or healthcare case management.
Experience integrating AI agents with Oracle Fusion, SAP S/4HANA, Salesforce or ServiceNow.
Hands-on experience with reinforcement learning, RLHF, RLAIF or Direct Preference Optimization.
Experience taking a fine-tuned or distilled LLM/SLM into production.
Experience implementing confidential computing or private AI environments.
Experience with NVIDIA NIM, NeMo, CUDA or GPU scheduling.
Experience designing AI platforms capable of operating across AWS, Azure, Google Cloud Platform and on-premises infrastructure.
FinOps experience specifically focused on GPU and LLM workloads.
Experience with multi-tenancy, usage metering, chargeback and token-level cost allocation.
Published research, patents or recognized open-source contributions related to agent infrastructure, LLM evaluation or AI security.
Microsoft Azure AI Engineer, AWS Machine Learning Specialty, Google Professional Machine Learning Engineer or equivalent certification.
TOGAF, Kubernetes CKA/CKAD or advanced cloud architecture certification.
Master's degree or Ph.D. in Computer Science, Artificial Intelligence, Machine Learning or a related discipline.
Candidate Validation Requirements:
Candidates must be prepared to describe at least two production implementations, including the business problem and measurable outcome; complete architecture; models and orchestration frameworks; RAG or GraphRAG design; agent planning, memory and tool use; evaluation and hallucination controls; security and Responsible AI; cloud deployment; monitoring; and latency, accuracy and cost results.
Candidates whose experience is limited to prototypes, demonstrations, copilots or proof-of-concept implementations will not be considered.