Site Reliability Engineer W2 Only F2F Interview

Atlanta, GA, US • Posted 6 hours ago • Updated 6 hours ago
Contract W2
12 Months
No Travel Required
On-site
$50 - $55/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Apache Kafka
  • Automated Testing
  • Backend Development
  • API
  • API Management
  • Analytical Skill
  • Dynatrace
  • Engineering Design
  • FOCUS
  • Database
  • Debugging
  • Design Patterns
  • Docker
  • Continuous Integration
  • Customer Facing
  • Dashboard
  • Grafana
  • IT Management
  • Incident Management
  • Continuous Delivery
  • Continuous Improvement
  • Computer Networking
  • Configuration Management
  • MEAN Stack
  • Management
  • Mentorship
  • Instrumentation
  • JPA
  • Java
  • Kubernetes
  • Cloud Computing
  • Collaboration
  • Communication
  • Budget
  • CHAOS
  • RabbitMQ
  • Resource Management
  • Performance Engineering
  • Production Support
  • Microsoft Azure
  • New Relic
  • NoSQL
  • Service Level
  • Software Development
  • Software Engineering
  • OAuth
  • SAFE
  • Scalability
  • Spring Security
  • Service Design
  • Splunk
  • Spring Framework
  • CPU
  • Systems Design
  • Terraform
  • Caching
  • Messaging
  • Microservices
  • Storage
  • Testing

Summary

SENIOR SITE RELIABILITY ENGINEER (SRE)

Java Microservices, Spring Boot & AKS Reliability Engineering

EMPLOYMENT TYPE

Full-time / Contract

SENIORITY

Senior

PRIMARY FOCUS

Microservices & AKS

 

Position Summary

We are seeking a Senior Site Reliability Engineer with strong hands-on software development experience in Java, Spring Boot, and cloud-native microservices. The ideal candidate is a backend software engineer who applies an SRE mindset to service design, production readiness, automation, observability, performance, and resiliency.

This role will focus on improving the reliability of distributed services deployed on Azure Kubernetes Service (AKS). The successful candidate must be comfortable reading and modifying application code, troubleshooting services across the application and Kubernetes layers, implementing telemetry, and engineering permanent solutions to production reliability risks.

ROLE EXPECTATION  This is not a production support, Kubernetes administration, monitoring-tool administration, or ticket-driven operations role. Recent hands-on Java/Spring Boot microservices development is required.

Primary Objectives

·   Engineer reliable, scalable, and observable Java/Spring Boot microservices running on AKS.

·   Improve service resiliency through code, architecture changes, automation, and production readiness standards.

·   Build proactive detection and diagnostics using logs, metrics, traces, and service-level indicators.

·   Reduce customer impact and operational toil by addressing reliability risks at the application and platform layers.

·   Partner with development and platform teams to improve deployment safety, runtime stability, and service ownership.

Key Responsibilities

Java & Spring Boot Microservices Engineering

·   Design, develop, review, and enhance Java-based microservices using Spring Boot and related frameworks.

·   Build secure, scalable, and well-defined REST APIs and event-driven services.

·   Apply software engineering practices including clean code, automated testing, code reviews, dependency management, and design patterns.

·   Review service architecture and application code to identify reliability, performance, scalability, and supportability gaps.

·   Troubleshoot complex failures across application logic, service dependencies, data stores, messaging components, and external integrations.

AKS & Cloud-Native Reliability

·   Deploy, operate, and troubleshoot microservices running on Azure Kubernetes Service (AKS).

·   Understand Kubernetes workloads, pods, deployments, services, ingress, ConfigMaps, secrets, health probes, resource limits, autoscaling, and rollout behavior.

·   Partner with platform engineering teams on AKS networking, identity, security, node capacity, workload placement, and cluster-level dependencies.

·   Analyze pod restarts, container failures, memory and CPU constraints, scaling behavior, readiness failures, and deployment issues.

·   Improve deployment strategies, rollback readiness, workload stability, and safe production releases.

Reliability & Resiliency Engineering

·   Define and manage SLIs, SLOs, Error Budgets, and service health indicators for critical microservices.

·   Implement and validate timeouts, retries, circuit breakers, bulkheads, rate limiting, caching, idempotency, and graceful degradation.

·   Conduct production readiness and dependency risk reviews for new services and major changes.

·   Participate in load, performance, capacity, and resiliency testing, including controlled dependency failure scenarios.

·   Identify systemic reliability risks and drive corrective application or platform engineering changes before customer impact occurs.

Observability & Diagnostics

·   Implement structured logging, application metrics, distributed tracing, and correlation identifiers using OpenTelemetry, Micrometer, or comparable frameworks.

·   Create actionable dashboards and alerts that reflect service health, dependency behavior, saturation, latency, throughput, and errors.

·   Correlate telemetry across APIs, microservices, messaging, databases, AKS workloads, and third-party systems.

·   Improve diagnostic detail in application code so failures can be identified quickly without relying on manual investigation.

·   Partner with SRE and engineering teams to reduce noisy alerts and improve proactive detection.

Software Delivery & Automation

·   Build and improve CI/CD pipelines for compile, test, scan, package, deploy, validate, and rollback activities.

·   Automate health validation, release verification, diagnostics, scaling checks, and repetitive operational tasks.

·   Build reusable libraries, templates, and standards for service instrumentation and reliability patterns.

·   Use Infrastructure as Code and configuration management practices to support consistent environments.

·   Mentor engineers on cloud-native design, application reliability, observability, and operational ownership.

Incident Management & Continuous Improvement

·   Provide hands-on technical leadership during incidents involving Java services, AKS workloads, integrations, or shared dependencies.

·   Use code, telemetry, Kubernetes events, deployment history, and dependency data to diagnose failures.

·   Drive blameless post-incident reviews and ensure corrective actions result in permanent engineering improvements.

·   Reduce mean time to detect, diagnose, recover, and prevent recurrence through automation and better system design.

Required Qualifications

Java & Microservices Development

·   7+ years of software engineering experience with significant recent hands-on backend development.

·   Strong Java development experience and deep practical knowledge of Spring Boot.

·   Experience designing, building, testing, and supporting production microservices and REST APIs.

·   Strong understanding of distributed systems, synchronous and asynchronous communication, failure handling, and service ownership.

·   Experience with automated unit, integration, contract, and component testing.

·   Ability to read, debug, modify, and review production application code.

AKS & Kubernetes

·   Hands-on experience deploying, operating, and troubleshooting applications on Azure Kubernetes Service (AKS).

·   Strong understanding of containers, Docker, Kubernetes workload resources, service discovery, ingress, probes, secrets, configuration, and autoscaling.

·   Experience diagnosing application failures using pod logs, Kubernetes events, resource metrics, and deployment status.

·   Understanding of Kubernetes networking, identity, security, resource management, and rollout strategies.

·   Experience working with Helm, Kubernetes manifests, GitOps, or comparable deployment methods.

Cloud, Integration & Data

·   Experience with Azure services such as AKS, API Management, Service Bus, Key Vault, Azure Monitor, databases, and storage.

·   Experience with messaging or event platforms such as Azure Service Bus, Kafka, RabbitMQ, or Event Hubs.

·   Experience with relational and/or NoSQL databases and practical knowledge of query performance, connection management, and failure handling.

·   Understanding of API security, OAuth 2.0, JWT, managed identity, secrets management, and certificate-based communication.

SRE & Observability

·   Strong understanding of SRE principles, including SLIs, SLOs, Error Budgets, production readiness, incident response, and toil reduction.

·   Hands-on experience with OpenTelemetry, structured logging, application metrics, distributed tracing, and alerting.

·   Experience with an observability platform such as Splunk, Datadog, Dynatrace, Grafana, Prometheus, New Relic, or equivalent.

·   Strong analytical, troubleshooting, communication, and cross-functional collaboration skills.

Preferred Qualifications

·   Experience with Java 17 or later and modern Spring ecosystem components.

·   Experience with Spring Cloud, Spring Security, Resilience4j, Micrometer, JPA, or comparable frameworks.

·   Experience with Istio, service mesh, KEDA, or advanced AKS scaling patterns.

·   Experience with Terraform, Bicep, or other Infrastructure as Code technologies.

·   Experience supporting high-volume digital commerce, ordering, payment, loyalty, or customer-facing platforms.

·   Experience with performance engineering, JVM tuning, chaos engineering, or fault injection.

·   Experience leading technical designs and mentoring software engineers.

Success Measures

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10123687
  • Position Id: 9094193
  • Posted 6 hours ago
Contact the job poster
Mary Menet

Mary Menet

Technical Recruiter @ Encore Consulting Services
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Atlanta, Georgia

•

Today

Full-time

USD 64,000.00 - 76,750.00 per year

Alpharetta, Georgia

•

3d ago

Easy Apply

Full-time

100000 - 105000

Alpharetta, Georgia

•

Yesterday

Easy Apply

Full-time

100,000 - 110,000

Atlanta, Georgia

•

Today

Full-time

USD 128,000.00 - 153,750.00 per year

Search all similar jobs