Sr AI Platform Engineer
Remote - ocassional oniste in Centreville, VA (Must be local to Centreville, VA)
No C2C or 3rd Party Candidates
About the Opportunity
We are seeking a Senior AI Platform Engineer to help build the secure, scalable infrastructure that enables enterprise adoption of AI. This is a hands-on engineering role focused initially on building a BYOK (Bring Your Own Key) foundation for coding agents, including GitHub Copilot and Claude Code, with governed model access through Azure AI Foundry and AWS Bedrock.
The platform will provide secure model access, AI gateways, MCP-based integrations, observability, governance, security controls, and AI FinOps capabilities.
Over time, the role will expand from coding-agent enablement into a broader enterprise AI platform supporting AI applications, agentic workflows, enterprise MCP integrations, and reusable agent-development capabilities.
What You Will Do
Build the AI Coding-Agent Foundation
Design and deploy secure, governed BYOK patterns for GitHub Copilot and Claude Code.
Route model requests through controlled Azure and AWS environments rather than unmanaged vendor defaults.
Administer and govern GitHub Copilot SaaS, including licensing, access policies, usage visibility, and configuration.
Integrate Azure AI Foundry and AWS Bedrock for approved model hosting and model routing.
Develop reusable infrastructure templates, repositories, deployment pipelines, onboarding processes, and implementation guides.
Support secure developer environments, private networking, identity federation, secrets management, and least-privilege access.
Build AI Gateway & MCP Gateway Capabilities
Design and operate AI Gateway / Azure API Management capabilities for model routing, authentication and authorization, model allow-lists, token budgets, rate limiting, circuit breakers, content filtering, and usage/cost attribution.
Build secure MCP Gateway capabilities allowing AI agents to access approved enterprise data, APIs, repositories, developer tools, ticketing systems, and SaaS services.
Establish governance processes for MCP servers, APIs, tools, and data connectors.
Implement trust-boundary controls covering data classification, tenant isolation, cross-cloud routing, and egress.
Security, Governance & Compliance
Implement controls for model eligibility, developer access, approved use cases, prompt/response handling, data protection, and tool-call authorization.
Partner with Security, Legal, Risk, Cloud, and Federal stakeholders.
Support security and compliance requirements involving NIST 800-53, NIST 800-171, CMMC, FedRAMP, FISMA, and CUI.
Design secure patterns for encryption, audit logging, authorized model endpoints, tenant selection, and controlled egress.
Implement human review, approval, exception, escalation, and rollback processes for higher-risk AI use cases.
Observability, AgentOps & AI FinOps
Build unified telemetry covering model calls, tool calls, usage, latency, errors, token consumption, cost, policy decisions, and user/team attribution.
Develop AI FinOps dashboards for token consumption, model/API usage, cost attribution, budget thresholds, anomaly detection, forecast vs. actual spend, adoption trends, and showback/chargeback.
Establish evaluation and quality processes, including regression testing, red-team scenarios, failure detection, and release gates.
Define SLOs, alerting, runbooks, incident response, root-cause analysis, and production-readiness standards.
Developer Enablement
Create onboarding documentation, sample projects, reference architectures, secure defaults, and model-selection guidance.
Establish reusable patterns for coding-agent workflows, secure prompt/data handling, repository access, code review, and tool usage.
Partner with Enterprise Architecture, Cloud Engineering, Security, SRE, Developer Productivity, Legal/Risk, and application engineering teams.
Help establish engineering standards for AI agents, BYOK, MCP, gateway policies, observability, and secure AI development.
Phase 2 – Enterprise AI Platform
Build reusable agent development platforms using Azure AI Foundry and AWS AgentCore.
Develop secure agent runtime and deployment patterns.
Provide governed model access through AI gateways.
Enable MCP-based access to enterprise systems and data.
Implement AI application testing and evaluation frameworks.
Develop model usage metering and cost-management capabilities.
Support tenant isolation and data classification.
Partner on platform roadmap development and enterprise AI standards.
Required Qualifications
8+ years of software, cloud, platform, DevOps, or infrastructure engineering experience.
3+ years building enterprise cloud or platform services.
1–2+ years supporting AI/ML, LLM, Generative AI, or agentic systems.
Strong hands-on Python development experience.
Infrastructure-as-code experience with Terraform, Bicep, CloudFormation, or equivalent.
Experience with containers, Kubernetes and/or serverless architectures.
Experience with CI/CD, secrets management, private networking, and enterprise identity.
Hands-on experience with Azure AI Foundry, Azure API Management, Azure Monitor / Application Insights, Microsoft Entra ID, Azure Key Vault, Azure Private Link, and Azure Policy.
Working knowledge of AWS Bedrock, AWS GovCloud, AWS IAM, AWS CloudWatch, and AWS networking.
Experience designing or operating AI/LLM gateways, model-routing layers, API gateways, MCP servers/gateways, tool/function calling, RAG pipelines, or agent orchestration platforms.
Experience administering or governing GitHub Copilot SaaS, including access management, policy configuration, licensing, and usage reporting.
Strong understanding of security architecture, including least privilege, RBAC/ABAC, Conditional Access, encryption, egress controls, audit logging, data classification, and secrets management.
Experience supporting regulated or compliance-driven environments.
Preferred Qualifications
Experience with GitHub Copilot Enterprise, GitHub Copilot BYOK, Claude Code, Anthropic SDKs, OpenAI/Azure OpenAI APIs, or AWS Bedrock.
Experience with MCP, A2A, LangGraph, Semantic Kernel, LlamaIndex, LangChain, or AWS AgentCore.
Experience implementing AI gateways with model allow-lists, content safety, prompt-injection defenses, usage metering, routing, rate limits, circuit breakers, and cost attribution.
Experience with Microsoft Purview, Defender for Cloud, Microsoft Sentinel/SIEM, OpenTelemetry, Langfuse, Arize, LangSmith, or similar AI observability/evaluation platforms.
Experience supporting Federal, Defense Industrial Base, GCC High, Azure Government, AWS GovCloud, or CUI environments.
Experience with AI FinOps, cloud cost management, token economics, cost allocation, forecasting, anomaly detection, and showback/chargeback.
Experience defining enterprise platform roadmaps, operating models, onboarding processes, SLOs, and operational runbooks.
Preferred Certifications
Microsoft Azure Solutions Architect Expert
Microsoft Azure AI Engineer Associate
Microsoft Azure Security Engineer Associate
AWS Solutions Architect
AWS Security Specialty
AWS Machine Learning
GitHub Copilot / GitHub Actions / GitHub Advanced Security
FinOps Certified Practitioner
Terraform or Kubernetes certification
CISSP, CCSP, Security+, or comparable certification
What Success Looks Like
A production-ready coding-agent BYOK foundation for GitHub Copilot and Claude Code.
Governed model access through Azure AI Foundry and AWS Bedrock.
An operational AI Gateway / Azure APIM platform.
An enterprise MCP Gateway with connector onboarding, access controls, audit logging, and decommissioning processes.
Embedded AI security, governance, compliance, and CUI handling controls.
AI observability and FinOps dashboards covering usage, tokens, costs, reliability, and policy decisions.
Repeatable developer onboarding processes that allow teams to adopt AI capabilities securely.
A roadmap for extending the platform to broader enterprise AI applications and agent-development platforms.
Why This Role Matters
This role creates the secure operating foundation for enterprise adoption of AI. Coding agents are the initial high-priority use case, while the long-term objective is a reusable enterprise AI platform supporting secure AI applications, MCP-enabled integrations, and governed agent development across Federal and Commercial environments.
Keywords / Skills
Senior AI Platform Engineer, AI Platform Engineer, AI Infrastructure, Generative AI, GenAI, LLM, LLMOps, Agentic AI, AI Agents, Coding Agents, GitHub Copilot, Copilot BYOK, Claude Code, Azure AI Foundry, AWS Bedrock, AWS AgentCore, Azure API Management, AI Gateway, MCP, MCP Gateway, Model Context Protocol, Python, Terraform, Bicep, Kubernetes, Docker, CI/CD, Azure, AWS, Azure Government, AWS GovCloud, Entra ID, Key Vault, Private Link, Azure Policy, IAM, API Gateway, RAG, LangGraph, LangChain, Semantic Kernel, LlamaIndex, OpenTelemetry, AI Observability, AgentOps, AI FinOps, Cloud FinOps, CMMC, FedRAMP, FISMA, NIST, CUI, DevSecOps, Cloud Security, Platform Engineering, Enterprise AI.
Please send resume to: