JOB DESCRIPTION:
- Build features end to end across the Nash platform — from the React/TypeScript front end through the Java/Spring Boot services to the Oracle schema behind them.
- Extend the asynchronous LLM execution pipeline. Work on the message-driven services that accept requests, queue them by priority, call the model provider, and return results — including retry, failover and error handling semantics.
- Integrate and maintain model provider connections across Azure OpenAI and AWS Bedrock: new model onboarding, routing, provider SDK upgrades, and the differences in how each provider reports usage and errors.
- Design and implement throughput controls — token- and request-per-minute rate limiting, backpressure, and circuit breaking — so the platform scales toward its 1M+ requests/day target without exhausting provider quota.
- Own usage, cost and token accounting. Extend the data model and reporting that tracks input, output and reasoning tokens per request, and surface it in the admin UI.
- Deliver the self-service authoring experience that lets business teams create and tune AI use cases without hands-on engineering support.
- Work the full change lifecycle in a regulated environment — design, code review, governance and compliance gates, and controlled promotion across environments.
- Build and maintain observability — structured logging, service metrics, and dashboards in Splunk, Datadog and Dynatrace; diagnose production issues across a distributed, queue-based system.
- Harden the platform — service authentication and authorisation, secrets handling, and dependency currency.
Collaborate with DBAs, platform, networking and compliance teams, whose sign-off the work depends on.
**Required**
- 5+ years building and running production web applications end to end
- **React 18 + TypeScript** — hooks, modern state management (Redux Toolkit / RTK Query), and experience with a high-volume data grid (ag-grid or equivalent)
- **Java 17 + Spring Boot** — REST APIs, Spring Data JPA, dependency injection, configuration management
- Strong **relational database and SQL** skills — schema design, query tuning, and working with DBA-applied migrations rather than ORM-generated DDL
- Solid grasp of **asynchronous and concurrent programming** — thread pools, executors, and the failure modes that come with them
- Git-based workflow with peer review; CI/CD pipelines
- Able to work within formal change control, and to write for an audience that includes non-engineers
**Strongly preferred**
- **Project Reactor / reactive Spring** (`Mono`, `Flux`, WebFlux) — core to the execution services
- **Message-oriented middleware** — IBM MQ, JMS, ActiveMQ or Kafka; consumer concurrency, redelivery, dead-letter handling
- **LLM API integration experience** — Azure OpenAI, Bedrock or OpenAI directly; prompt handling, token accounting, rate limits, and 429/backoff semantics
- **Kubernetes and Docker** in an enterprise setting
- **Micro-frontend architecture** — Webpack Module Federation or similar
- Oracle specifically, including PL/SQL
- Observability tooling: Splunk, Datadog, Dynatrace, PrometheMicrometer
**Nice to have**
- Azure API Management, or API gateway policy work generally
- Financial services or another regulated industry
- Performance and load testing at scale
- Maven and private artifact repositories (Artifactory)
Education:
Bachelor's degree in Computer Science, Software Engineering, Information Systems, or a related technical field — or an equivalent combination of education, training and experience.