Role: Senior Platform Engineer (AI Gateway, Identity & Observability)
Location: Oaks, PA
Type : Long Term Contract
Exp. Required: 10+ years
Need candidate who can work onsite from Day 1 (Hybrid basis)
Experience: 10+ years overall, including 5+ years hands-on API platform, identity or observability engineering
RESPONSIBILITIES
The platform is being designed as the engagement progresses, so responsibilities will evolve. Indicative, and not limited to:
- Deliver the AI gateway that routes model traffic across cloud-hosted models, other providers and on-premises inference, presenting one consistent, provider-agnostic interface to every consuming agent.
- Implement gateway policy: routing and backend pools, token-based rate limiting and quotas, retries, failover and circuit breaking, caching, and request and response transformation.
- Harden an existing service for production and run a controlled migration to the long-term gateway, moving consumers across without changing the contract they already integrated against.
- Build the agent identity and registration path: workload or service identity for agents, registration on deployment, and the promotion gate that governs what reaches production. Full automation is not assumed; part of the work is automating what can be automated and designing a clean, documented path for the steps that require human approval.
Skills Required
- 5+ years hands-on engineering across API platforms, identity or observability, with production depth in at least one and working competence in the others. All three areas are in scope for this position.
- 4+ years building and operating API gateways in production: routing, policy authoring, rate limiting and quotas, authentication, transformation and versioning. Azure API Management is strongly preferred; comparable depth in Apigee, Kong or an equivalent enterprise gateway is acceptable.
- 3+ years enterprise identity and access engineering: OAuth2, OpenID Connect and JWT validation, service principals and workload or managed identity, and secrets management using an enterprise vault.
- 3+ years hands-on OpenTelemetry or equivalent distributed tracing: instrumentation, collectors, resource attributes, span design, and trace, metric and log pipelines.
- 3+ years in a site reliability, production engineering or platform operations role, or equivalent hands-on responsibility for a service other teams depend on: service level objectives and error budgets, alerting design, on-call and incident response, and post-incident analysis. This position owns whether the telemetry is trustworthy, not only whether it is being collected.
- 2+ years working with LLM or AI workloads in production, including model routing across providers, token-based limits and quotas, streaming responses, content filtering, and how inference cost accrues and is attributed.
- Experience with LLM observability tooling, such as Langfuse, LangSmith, Arize or an equivalent platform, including how agent traces differ from conventional application traces.
- 3+ years production API operations: failover, circuit breaking, caching, load testing to prove capacity, and incident response for a service other teams depend on.
- 2+ years integrating third-party or self-hosted services into an enterprise network, including private connectivity or controlled egress, DNS resolution, TLS and certificate management, and working the firewall and security review needed to get each path approved.
- Working knowledge of cloud infrastructure on at least one major hyperscaler: networking, identity, container platforms and infrastructure-as-code.
- Programming competence in Python, C# or an equivalent language, sufficient to build and maintain policy extensions, callout services and instrumentation libraries.
PREFERRED
- Familiarity with Azure AI Foundry observability and Azure Monitor or Application Insights, including agent tracing, continuous evaluation and how a managed observability plane compares with a self-hosted one such as Langfuse. This platform may run one, the other, or both, and the choice is still open.
- Azure API Management depth, including policy expressions, reusable policy fragments and the AI or LLM policy set.
- Microsoft Entra ID experience, including managed identity, app registrations and conditional access.
- Experience with reliability engineering for AI or machine learning systems, including quality regression detection, drift monitoring or evaluation-driven alerting.
- Familiarity with emerging agent identity models and how agent-to-service authentication differs from conventional service-to-service patterns.
- Experience implementing chargeback or showback for a shared platform service.
- Content safety or guardrail services, such as Azure AI Content Safety, or custom filtering and moderation services.
- Enterprise workflow integration, such as ServiceNow APIs for approval-driven provisioning.
- Experience in financial services or another regulated industry, where security review governs the pace of change.
- Familiarity with agent frameworks and how agents consume model endpoints, tools and memory.
Tekshapers is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.