Summary
Build and scale a production multi-agent AI platform serving thousands of internal users across multiple business units. Monthly release cadence, real users, real latency, real cost.
What You'll Own
LLM-driven orchestrator that routes user intent across a portfolio of specialized agents delegation, memory, response validation, capability discovery.
Agent selection layer hybrid retrieval (vector RAG over a capability registry) plus closed-set LLM selection with JSON-schema-constrained outputs.
Multi-agent SDK / gateway FastAPI service hosting many agents behind path-prefix routing, per-agent tool registries, session-scoped conversational context.
Tool-driven agents 15 30 tools per agent composed dynamically by an LLM; owns tool contracts, guardrails, and evaluation.
Data API layer parameterized endpoints between agents and databases; LLMs never touch DBs directly.
Partner-team onboarding versioned A2A contract, bring-your-own-agent registration, auto re-embedding.
Core AI Engineering
Production LLM systems: RAG, tool/function-calling loops, structured outputs, hallucination guards, closed-set selection.
Multi-agent orchestration: A2A protocols, session affinity, human-in-the-loop gating, kill switches, graceful degradation.
Vector search + embeddings at scale (sub-second retrieval over thousands of docs).
Evaluation & safety: PII/PHI masking, audit trails, feedback-loop instrumentation, offline + online eval.
Platform / Infrastructure
Python 3.11+, FastAPI, async I/O, Pydantic.
Modern LLM stacks (Gemini, GPT, Claude) and agent frameworks (LangGraph, Agent SDKs).
Cloud (Google Cloud Platform or AWS): Kubernetes, object storage, workflow orchestration, Vertex/Bedrock-class services.
Redis, MongoDB, Oracle/Postgres, SSO + RBAC.
Observability: Prometheus, structured JSON logs, per-decision audit trails, p95 latency SLOs in seconds.
Skills: Digital : Python~Digital : Machine Learning~Digital : Artificial Intelligence(AI)~Generative AI
Experience Required: 4-6