We are seeking a Senior Platform Engineer to own US-hours production support and platform engineering for a business-critical product catalog application. The platform runs on AWS and spans an Apigee API gateway, containerised workloads managed through Rancher and Kubernetes, and a set of C# and .NET services backed by a relational database.
The application recently completed a migration to AWS, and the immediate priority is stabilisation: diagnosing and resolving connection, latency and reliability issues across the stack, then converting those findings into durable fixes rather than recurring workarounds. You will be the senior technical presence during US business hours, partnering with the offshore engineering team that delivered the migration and with the business stakeholders who depend on the platform.
Over time the role broadens beyond incident response. The platform's business logic is governed by a rules engine with SQL-backed rules storage, and you will build proficiency in that model so you can take on non-production rules work and reduce the load currently carried by a small number of senior support engineers.
• Provide US-hours production support for the product catalog platform, with primary focus on connection failures, latency, and general performance degradation.
• Triage and troubleshoot incidents end to end across the AWS infrastructure, Apigee gateway, Rancher and Kubernetes layer, C# services, and database tier — owning the issue rather than routing it onward.
• Investigate and resolve API-level failures including timeouts, routing and connection errors, policy misconfiguration, and gaps in logging and observability.
• Partner with the offshore team that delivered the AWS migration to stabilise and optimise the current environment, ensuring fixes are understood and retained on both sides.
• Work directly with platform stakeholders to establish root cause and drive sustainable remediation, distinguishing genuine fixes from temporary mitigations.
• Support, debug and optimise the C# codebase powering the platform and its API integrations.
• Develop working proficiency in the platform's rules engine and SQL-based rules storage, progressively taking on non-production rules work to offload senior support engineers.
• Contribute to documentation and structured knowledge sharing to reduce single-point-of-failure risk across the team.
AWS
• Hands-on experience troubleshooting production workloads in AWS.
• Working familiarity with the core services underpinning an enterprise application — compute, database, networking, load balancing and monitoring.
Apigee / API Gateway
• Production experience with Apigee or Apigee X, or a comparable enterprise API gateway.
• Demonstrated ability to debug API failures, timeouts, policy issues and routing problems under live conditions.
Kubernetes / Rancher
• Strong working knowledge of Kubernetes fundamentals — pods, deployments, services, and scaling behaviour.
• Experience managing clusters and workloads through Rancher or comparable tooling.
• Ability to diagnose pod instability, restart loops, resource constraints and related performance issues.
C# / .NET Application Services
• Proficiency in C# and .NET, particularly within API and service-oriented architectures.
• Strong debugging and troubleshooting capability in distributed systems, including the ability to trace a failure across service boundaries.
General
• Proven track record providing production support for complex, multi-tier applications.
• Strong analytical and problem-solving skills, with genuine comfort owning an issue from first report to closure.
• Clear communication with both engineering teams and business stakeholders during active incidents.
• Experience with rules engines or rule-based business logic in a production system.
• Strong SQL skills for performance analysis and data-level troubleshooting.
• Aptitude and appetite to learn a proprietary rules model and grow into rules work alongside the existing senior engineering team.
• Experience supporting an application through or immediately following a cloud migration.