Senior AI Platform Engineer - Nearshore

Hybrid in Toronto, ON, CA • Posted 1 day ago • Updated 1 day ago
Contract Independent
Contract W2
12 Months
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Artificial Intelligence
  • LangChain
  • Kubernetes

Summary

Position Title: Senior Platform Engineer - Nearshore

Location: Canada (Remote)

Type: Contract

  • Build, deploy, and operate the shared AI platform services that application teams depend on LLM gateway and proxy layers, model routing and fallback, MCP servers and registries, agent runtimes, and the supporting APIs around them with telemetry, health signals, and cost attribution designed in rather than added later.

  • Onboard and enable new model providers and model versions across environments, including routing rules, rate limits and quotas, model tiering and cost-aware selection, fallback behavior, and version deprecation paths, and make each provider's usage, latency, error, and spend profile visible and comparable.
  • Design and operate the MCP layer that exposes enterprise systems as tools to agents server deployment, tool registration and discovery, connection and session handling, safe tool-permission boundaries, and tracing of tool calls, retries, and failures.

  • Support the agent lifecycle on the platform publishing and versioning, registry and discovery, memory and session management, scaling and resilience and instrument agent steps, state transitions, and evaluation outcomes so behavior can be explained after the fact.

  • Run platform services on Kubernetes using GitOps practices: Helm and Kustomize manifests, declarative delivery, autoscaling, resource tuning, health checks, and zero-downtime or progressive rollout.
  • Implement authentication, authorization, and tenancy for the platform OAuth2/OIDC and SSO integration, JWT issuance and validation, role and group based access, virtual key and API key management, and secrets handling through a managed key vault.

  • Apply guardrails and policy enforcement, including PII detection with masking or blocking, prompt and response filtering, content policy, audit logging, and retention rules for prompts, completions, and traces.
  • Partner with engineering and developer teams to define observability standards for AI applications, LLM integrations, and agentic workflows, and make those standards the default path rather than an extra step.
  • Implement and maintain telemetry patterns in Langfuse, including traces, spans, prompts, completions, feedback, evaluation metadata, latency, errors, token usage, and model/provider context.
  • Instrument the platform itself with OpenTelemetry, metrics, and structured logs so gateway, MCP, and agent-runtime behavior is traceable end to end alongside application-level AI traces.
  • Build dashboards, reports, and alerts that help teams understand AI reliability, performance, quality, evaluation outcomes, and production behavior.
  • Design cost and usage observability across LLM vendors such as OpenAI, Anthropic, Google Gemini, and other providers, including attribution by application, team, user, model, workflow, and environment.
  • Create showback or chargeback-ready metrics for token usage, request volume, model mix, latency, cache behavior, evaluation runs, and vendor spend, and feed what they show back into routing, tiering, and capacity decisions.

  • Support AI evaluation practices by helping teams define test sets, golden datasets, scoring strategies, prompt and version comparisons, regression checks, and release readiness signals.
  • Analyze execution patterns across agent tooling such as Claude Code, in-house agent frameworks, LangChain, LangGraph, CrewAI, and Google ADK, and turn what you find into platform fixes and guidance.
  • Support retrieval and context-grounding infrastructure embedding models, vector and graph stores, ingestion and refresh pipelines, and retrieval quality measurement and tuning.
  • Build the developer-facing self-service surface: onboarding automation, provisioning workflows, templates, and internal portals or CLIs that let teams get access, register agents and tools, and ship without manual tickets.
  • Extend observability into the software delivery layer where relevant AI-assisted development telemetry, pipeline and pull-request lifecycle metrics, and DORA-style delivery signal

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10496476a
  • Position Id: 9102309
  • Posted 1 day ago
Contact the job poster
Sunil Joshi

Sunil Joshi

Recruiter @ Diligente Technologies
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

•

17d ago

Easy Apply

Contract

Depends on Experience

Remote

•

Yesterday

Easy Apply

Contract

Depends on Experience

Remote

•

Today

Easy Apply

Contract

90 - 110

Remote

•

3d ago

Easy Apply

Full-time

133000 - 164000

Search all similar jobs