Sr AI Platform Engineer
No Corp to Corp, 3rd Party or IC
Remote - Must travel to job site occassionally
Sr AI Platform Engineer
Remote - ocassional oniste in Centreville, VA (Must be local to Centreville, VA)
No C2C or 3rd Party Candidates
Position Overview
We are seeking a highly experienced Senior AI Platform Engineer to design, build, and operate the foundational infrastructure required to securely deploy, govern, monitor, and scale enterprise AI capabilities.
The initial 3–6 month focus will be on establishing a secure enterprise foundation for AI coding agents, including GitHub Copilot and Claude Code, using Bring Your Own Key (BYOK) deployment patterns. Model access will be routed through controlled cloud environments utilizing Azure AI Foundry and AWS Bedrock.
Following the initial implementation, this role will expand into a broader enterprise AI platform supporting governed model access, Model Context Protocol (MCP) integrations, secure agent runtimes, AI gateways, observability, AI FinOps, and reusable agent-development platforms.
This is a hands-on engineering role requiring strong experience across cloud infrastructure, DevOps, AI/ML platforms, security, governance, and enterprise architecture.
Key Responsibilities
AI Coding-Agent Platform
- Design and implement secure, reusable BYOK infrastructure patterns for GitHub Copilot and Claude Code.
- Route model requests through company-controlled cloud environments rather than unmanaged vendor defaults.
- Administer and govern GitHub Copilot SaaS environments, including licensing, access policies, configuration, usage visibility, and coordination with GitHub/Microsoft 365 administrators.
- Integrate Azure AI Foundry and AWS Bedrock for approved model hosting, routing, access policies, and secure connectivity.
- Develop infrastructure templates, sample repositories, CI/CD pipelines, and onboarding documentation.
- Design secure developer environments incorporating private networking, identity federation, secrets management, and least-privilege access.
- Build reusable platform capabilities that can support future enterprise AI applications.
AI Gateway & MCP Gateway
- Design and operate an AI Gateway / Azure API Management (APIM) capability supporting:
- Model routing
- Authentication and authorization
- Model allow-lists
- Token budgets
- Rate limiting
- Circuit breakers
- Content filtering
- Usage attribution
- Build an enterprise MCP Gateway that securely brokers access between AI agents and enterprise data, APIs, repositories, developer tools, ticketing systems, and approved SaaS platforms.
- Establish governance processes for MCP servers, APIs, tools, and data connectors.
- Implement connector onboarding, registration, approval, versioning, monitoring, and decommissioning processes.
- Enforce trust boundaries around data classification, tenant isolation, cross-cloud routing, and egress.
Security, Governance & Compliance
- Implement technical controls governing model eligibility, developer access, approved use cases, prompt/response handling, data loss prevention, and tool authorization.
- Partner with Security, Legal, Risk, Enterprise Architecture, and Federal stakeholders to establish appropriate controls.
- Support compliance requirements associated with NIST 800-53, NIST 800-171, CMMC, FedRAMP, FISMA, and DoD impact levels, where applicable.
- Design secure patterns for handling Controlled Unclassified Information (CUI).
- Implement appropriate tenant selection, encryption, authorized model endpoints, audit logging, and controlled egress.
- Establish human-review, approval, exception, escalation, and rollback processes for higher-risk AI use cases.
Observability, AgentOps & AI FinOps
- Develop unified telemetry for AI platforms, including model calls, tool calls, data access, latency, errors, token consumption, cost, policy decisions, and team/user attribution.
- Establish dashboards and reporting for AI usage and financial management.
- Track:
- Token consumption
- Model and API usage
- Cost attribution
- Budget thresholds
- Anomalies
- Adoption trends
- Forecast versus actual spend
- Showback/chargeback
- Develop AI evaluation and quality-control processes, including test harnesses, regression testing, red-team scenarios, hallucination/failure detection, and release gates.
- Establish production SLOs, monitoring, alerting, runbooks, incident response, root-cause analysis, and continuous platform hardening.
Developer Enablement
- Create developer onboarding guides, sample applications, approved model-selection guidance, and training materials.
- Develop reusable reference architectures for:
- Coding-agent workflows
- Secure prompt and data handling
- Repository access
- Pull request/code review workflows
- Tool-use guardrails
- Partner with Enterprise Architecture, Cloud Engineering, Security, SRE, Developer Productivity, Legal/Risk, and application engineering teams.
- Establish engineering standards for AI agents, BYOK, MCP, gateway policies, observability, and secure AI development.
Enterprise AI Platform Expansion
As the platform matures, expand capabilities beyond coding agents to support enterprise AI applications and agent platforms.
- Build and operate reusable agent-development platforms using Azure AI Foundry and AWS AgentCore.
- Establish secure templates, deployment pipelines, runtime patterns, testing/evaluation workflows, and operating standards.
- Provide governed model access through centralized gateways and policy enforcement.
- Enable enterprise applications and agents to securely access approved systems through MCP.
- Establish model, tenant, data-classification, security, usage-metering, and cost-management controls.
- Help define and execute the long-term enterprise AI platform roadmap.
Required Qualifications
- 8+ years of experience in software engineering, cloud engineering, platform engineering, DevOps, or infrastructure engineering.
- At least 3+ years of enterprise cloud/platform engineering experience.
- At least 1–2+ years of experience supporting AI/ML, LLM, generative AI, or agentic systems.
- Strong hands-on Python development experience.
- Strong infrastructure-as-code experience with Terraform, Bicep, CloudFormation, or equivalent.
- Experience with containers, Kubernetes and/or serverless architectures.
- Strong understanding of CI/CD, secrets management, private networking, and enterprise identity.
- Hands-on experience with:
- Azure AI Foundry
- Azure API Management
- Azure Monitor / Application Insights
- Microsoft Entra ID
- Azure Key Vault
- Azure Private Link
- Azure Policy
- Working knowledge of:
- AWS Bedrock
- AWS GovCloud
- AWS IAM
- CloudWatch
- AWS networking
- Multi-cloud/cross-cloud architectures
- Experience designing or operating LLM gateways, model-routing layers, API gateways, MCP servers/gateways, tool/function calling, RAG pipelines, or agent orchestration platforms.
- Experience administering or governing GitHub Copilot SaaS, including access management, licensing, policy configuration, and usage reporting.
- Strong security architecture knowledge, including:
- Least privilege
- RBAC/ABAC
- Conditional access
- Encryption
- Egress controls
- Audit logging
- Data classification
- Secrets management
- Understanding of regulated cloud environments and frameworks such as NIST, CMMC, FedRAMP, FISMA, and CUI requirements.
- Experience implementing observability for distributed systems, including logs, metrics, traces, dashboards, alerts, SLOs, incident response, and operational runbooks.
- Excellent communication skills with the ability to explain technical architecture, security risk, cost, compliance, and operational trade-offs to both technical and non-technical stakeholders.
Preferred Qualifications
- Experience with GitHub Copilot Enterprise, GitHub Copilot BYOK, Claude Code, Anthropic SDKs, OpenAI/Azure OpenAI APIs, or AWS Bedrock.
- Experience with MCP, A2A, LangGraph, Semantic Kernel, LlamaIndex, LangChain, AWS AgentCore, or similar AI orchestration frameworks.
- Experience developing AI gateways with:
- Model allow-lists
- Content safety
- Prompt-injection defenses
- Usage metering
- Model routing
- Rate limiting
- Circuit breakers
- Cost attribution
- Experience with Microsoft Purview, Defender for Cloud, Microsoft Sentinel/SIEM, policy-as-code, OpenTelemetry, Langfuse, Arize, LangSmith, or similar AI observability/evaluation technologies.
- Experience supporting the Defense Industrial Base (DIB), Federal environments, GCC High, Azure Government, AWS GovCloud, or CUI workloads.
- Experience implementing cloud or AI showback/chargeback models.
- Experience with AI FinOps, cloud cost management, token economics, cost allocation, budget controls, forecasting, anomaly detection, and financial reporting.
- Experience owning enterprise platform roadmaps, operating models, onboarding processes, SLOs, runbooks, and stakeholder communications.
Preferred Certifications
- Microsoft Certified: Azure Solutions Architect Expert
- Microsoft Azure AI Engineer Associate
- Microsoft Azure Security Engineer Associate
- AWS Certified Solutions Architect
- AWS Security Specialty
- AWS Machine Learning certification
- GitHub Copilot, GitHub Actions, or GitHub Advanced Security certification
- FinOps Certified Practitioner
- Terraform or Kubernetes certification
- CISSP
- CCSP
- CompTIA Security+
- Comparable cloud, AI, security, platform, or FinOps certifications
First-Year Success Measures
Within the first 12 months, the successful candidate will be expected to establish measurable enterprise capabilities including:
Coding-Agent BYOK Foundation
- Approved BYOK reference architectures for GitHub Copilot and Claude Code.
- Secure model access through Azure AI Foundry and AWS Bedrock.
AI Gateway
- Operational AI Gateway/APIM enforcing model access, routing, token budgets, logging, cost attribution, and policy controls.
MCP Gateway
- Enterprise MCP access patterns implemented for approved systems with connector onboarding, access controls, audit logging, and lifecycle management.
Governance & Compliance
- Operational model-access policies, data-classification controls, CUI handling patterns, audit trails, security evidence, and exception workflows.
Observability & AI FinOps
- Common dashboards providing visibility into adoption, token usage, model/API consumption, cost attribution, forecasts, anomalies, budget thresholds, policy exceptions, reliability, and operational health.
Developer Enablement
- Repeatable onboarding patterns that allow development teams to adopt approved AI capabilities without bypassing security, architecture, or compliance requirements.
Enterprise AI Platform
- Expansion of the coding-agent foundation into reusable model-access, MCP, and agent-development capabilities supporting broader enterprise AI applications.
Ideal Candidate
The ideal candidate is a hands-on senior/staff-level platform engineer who can bridge the gap between AI engineering, cloud infrastructure, cybersecurity, DevOps, and enterprise governance.
You should be comfortable moving from architecture and design into implementation, automation, troubleshooting, monitoring, and operational support. You will be expected to build secure, reusable platform capabilities rather than one-off solutions and to work effectively with engineering, security, architecture, compliance, and business stakeholders.
This position is particularly well suited for an engineer who has experience building enterprise AI platforms in regulated or highly secure environments and understands how to balance innovation, security, governance, developer productivity, reliability, and cost.