About the Role
We are looking for a Platform Product Engineer to be the product owner of a core engineering automation product part of our Internal Developer Platform (IDP). Your domain is platform orchestration, infrastructure automation, and end-to-end product deployment for single- and multi-tenant SaaS: the paved road that takes a change from commit to a healthy, observable production tenant. You are accountable for whether that paved road gets used and moves the numbers.
The role is a hybrid of product ownership and hands-on engineering. You direct one or more small developer pods that build the production-grade automation. Your own work covers scoping the problem, designing the solution, mapping the product capabilities, and building an end-to-end prototype that proves the approach. The pod then turns that prototype and capability map into hardened, operable production software. You work at the prototype and architecture level; the pod owns production implementation and operations.
Key Responsibilities
As a Platform Product Engineer, you will:
Own the Product — Scope, Solution, Capability Map, Prototype
- Own scope and solution — turn ambiguous platform problems into a clear scope, a designed solution, and a mapped set of product capabilities the org can rally around. You own the what and why of a core engineering product.
- Map product capabilities — define the capabilities the platform offers (provisioning, deployment, tenancy, data operations, Day-2), how they fit together, and how they sequence on the roadmap. Own the API contracts and the developer experience at the design level.
- Build the end-to-end prototype — prove the capability yourself with a working, runnable prototype: real self-service API, a real workflow behind it, and enough of a surface (portal or CLI) that a developer can see it work. The prototype validates the idea and becomes the spec the pod builds against. It is throwaway-to-hardened, not production-grade — thatʼs the podʼs job.
- Hand off to production, then stay accountable — the developer pod turns your prototype and capability map into production-grade automation. You stay accountable for the product outcome: that it ships, gets adopted, and moves the metrics.
Enable One or More Developer Pods
- Set direction for the pod(s) — bring the scope, solution, capability map, and prototype the pod needs to build production-grade automation with confidence. Break work down, sequence it, and keep the pod unblocked.
- Be the technical counterpart, not a hands-off PM — you review designs and PRs at the level of someone who could have built it, make pragmatic trade-offs between speed, safety, and maintainability with the pod, and raise the bar on API design, workflow durability, and operability.
- Own the backlog and the roadmap — a well-shaped backlog the pod can pick up and run, sequenced by dependency and risk, tied to the capabilities and outcomes you’re accountable for.
Treat Internal Engineers as Customers
- Build relationships with the developers you serve — talk to them informally and often, respond quickly, and use open-ended questions to tell a genuine problem from a nice-to-have. Document what you hear so the pod can act on it.
- Drive adoption, not just delivery — define the DevEx and reliability metrics the product should move, measure them honestly, and treat low adoption as a product bug you own.
- Reduce cognitive load and friction — shape self-service surfaces that abstract complex underlying systems behind safe defaults and strong guardrails, so engineers can provision, deploy, and operate autonomously from local dev through production.
Discover, Validate, Iterate
- Run discovery — interviews, jobs-to-be-done, and personas to find the problems worth solving; turn what you hear into a prioritized backlog.
- Prototype to validate, then descope — bias toward a thin, real, working prototype over an elaborate mock or an exhaustive research doc. Get it in front of the people whoʼd use it, learn, and cut scope hard so the pod builds the validated core first.
- Verify real usage and close the loop — instrument what ships, watch how (and whether) it gets adopted, feed that back into the roadmap, and tell users what you built and why when you decline a request. Treat the product as living, not delivered.
Cloud Operations Outcomes & Deployment Value Stream
- Own the value stream, define the metrics — instrument and track flow from idea to production tenant and value stream of the process being automated.
- Shape AI-assisted operations where it cuts toil — safe, auditable, human-in-the-loop workflows (alert enrichment, runbook-assisted diagnosis, onboarding assistants), with human review on anything that changes production.
Required Qualifications
Candidates must have sufficient depth in each of the following to design the solution, build an end-to-end prototype, and technically direct a pod through the production build.
- 6+ years in software or platform engineering, including hands-on experience shipping production infrastructure or platform systems, with demonstrated technical credibility and a clear motivation to move from a pure implementation track to a product-owning role.
- Deep cloud and cloud-native experience in at least one hyperscaler (AWS, Google Cloud Platform, or Azure; AWS preferred) across the core building blocks of a platform: Kubernetes (EKS/GKE/AKS — resource model, scheduling, ingress, autoscaling, cluster operations) and managed compute; landing zones (Organizations / Control Tower / LZA or equivalent) and guardrails; networking (VPCs, routing, security groups, load balancers, DNS, private connectivity); managed data (RDS/Aurora/Cloud SQL plus caching, queues, or NoSQL, with HA and backup); storage (object and block/file); and IAM and security (roles, workload identity, secrets, least-privilege), together with cloud-native architecture and sound operational practices.
- Self-service API architecture — experience designing and building the API surface through which engineers and automation provision and operate infrastructure: RESTful and/or gRPC APIs, versioning, error handling, and stable consumer contracts.
- API-based automation and durable workflow orchestration — experience with API-driven automation and long-running, workflow-driven processes (Temporal, Cadence, Step Functions, Argo Workflows, or a control-plane reconciliation model) using retries, idempotency, saga/compensation, and resume-after-failure to drive end-to-end provisioning, data operations, or platform work at scale, with a solid grasp of distributed-systems fundamentals (eventual consistency, failure handling, resilience).
- Platform automation, IaC, and config-as-code — experience automating infrastructure and its configuration declaratively with IaC (Terraform, Pulumi, or similar) and config-as-code (Ansible, Helm, Kustomize, or similar).
- End-to-end deployment for single- and multi-tenant SaaS — experience with the delivery path from commit to production tenant: environment promotion, GitOps/CD (ArgoCD, Flux, or similar), progressive/blue-green/canary rollout and rollback, and tenant provisioning, isolation, and configuration for both single- and multi-tenant models.
- Deployment value stream and cloud-ops metrics — experience defining and instrumenting DORA metrics (deployment frequency, lead time, change-failure rate, time-to-restore), lead/cycle time, and cloud cost, and using them to drive delivery.
- Hands-on prototyping ability — able to build a working, end-to-end prototype (API, workflow, and a thin surface) in Go and/or Python, with TypeScript/React for the surface, to validate an approach before the pod invests in the production build.
- Product ownership on platform or developer-tooling work — discovery, roadmap and backlog ownership, defining success metrics, iterating on real adoption, and treating the engineers served as customers; a track record of moving a platform capabilityʼs adoption, not only shipping it.
- Experience directing engineers — leading a small team or pod as a tech lead or product owner: setting scope, reviewing designs and code, sequencing work, and acting as the technical counterpart to the engineers.
Preferred Qualifications
- Prior experience as a product engineer, platform tech lead, or technical product owner for an engineering/infrastructure product — and the ability to articulate how you split your own hands-on work from the podʼs.
- Product analytics fluency — instrumenting features and building dashboards (PostHog, Amplitude, Mixpanel, Looker, or similar) to drive roadmap decisions.
- Experience with internal developer platform (IDP) patterns, platform-as-a-product methodology, or Backstage-style developer portals.
- Value-stream management or DevOps-metrics tooling (DORA/flow dashboards) — instrumenting and reporting delivery flow, not just building it.
- Kubernetes operator / controller development (controller-runtime, CRDs) or extending other orchestration substrates beyond containers.
- Multi-tenant SaaS at scale — pooled vs. siloed tenancy trade-offs, per-tenant quota/rate governance, noisy-neighbor isolation.
- FinOps / cloud cost governance and policy-as-code (OPA/Gatekeeper, Kyverno, or similar) for guardrails on the paved road.
- Frontend fluency (React / Next.js, TypeScript) for building richer prototype surfaces yourself.
- PostgreSQL and event-driven architectures (Kafka, NATS, or similar).
- Working knowledge of an AI/agent integration pattern (LLM APIs, RAG, agents, evals) applied to operational or DevEx tooling.
- Contributions to open-source infrastructure or developer-tooling projects; prior time at infrastructure, developer-tools, or platform-engineering companies.
First 30 / 60 / 90 Days:
First 30 days — Build context, meet the engineers you serve as customers and the developer pod(s) you’ll direct, and get hands-on with what we have today. Understand how the current workflows run and where the friction lives — from the people who feel it, not just the docs.
First 60 days — Own one platform capability: scope it, design the solution, map the capability, and build the end-to-end prototype that proves it. Hand it to the pod with a shaped backlog, and define the adoption and value-stream metrics it should move.
First 90 days — Take a cross-team capability from problem to measurable production adoption — for example, self-service tenant provisioning, environment lifecycle automation, or the deployment-orchestration foundations behind end-to-end SaaS delivery. You prototype and direct; the pod productionizes. Stand up the value-stream instrumentation (DORA / flow metrics) that proves it moved the numbers, close the loop with your users, and establish a durable pattern for how the pod ships platform-as-product.