Job Title: Site Reliability Engineer (SRE) – Production Services
Location: Pittsburgh, PA (Onsite)
Client Address: 500 Grant Street, Pittsburgh, PA 15219
Employment Type: Contract
Rate: $55–58/hr
Experience Required: 8–10 Years
Note: Local candidates only.
We are seeking an experienced Site Reliability Engineer (SRE) to join our Production Services team. This role is ideal for professionals with strong expertise in Java Spring Boot, Apache Kafka, DevOps, and CI/CD automation who are passionate about building resilient, highly available, and self-healing production platforms.
The ideal candidate will drive automation, improve operational efficiency, implement observability solutions, and enhance production reliability through modern SRE and DevOps practices.
Top Required Skills:
Key Responsibilities:
Automate high-volume production support requests and operational workflows to improve efficiency and reduce manual effort.
Develop self-service capabilities and automated remediation solutions for recurring operational tasks.
Design resilient operational workflows with auditability, consistency, and fault tolerance.
Implement automated retry mechanisms and intelligent backoff strategies for recurring production failures.
Define, implement, and manage Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical applications and batch processes.
Apply error budget principles to support release management and reliability improvements.
Improve batch processing reliability through standardized recovery patterns and proactive monitoring.
Build observability dashboards to monitor incidents, failure rates, repeat issues, automation coverage, and production health.
Enhance operational reporting across incidents, changes, and problem management processes.
Develop comprehensive runbooks and convert manual operational procedures into automated workflows.
Drive self-service capabilities for common operational requests and recurring production activities.
Implement self-healing capabilities to automatically detect and remediate production issues.
Optimize monitoring and alerting platforms, including Moogsoft, to reduce alert fatigue and improve signal quality.
Leverage automation and AI-driven operational practices to proactively resolve recurring production issues.
Collaborate with Development, Infrastructure, DevOps, and Operations teams to continuously improve platform reliability.
Required Qualifications:
8–10 years of experience in Site Reliability Engineering, Production Support, DevOps, or Platform Engineering.
Strong hands-on experience with Java and Spring Boot.
Strong expertise in Apache Kafka and event-driven architectures.
Experience with DevOps practices and CI/CD automation.
Experience building automation solutions using scripting and infrastructure automation tools.
Strong knowledge of production monitoring, observability, logging, and incident management.
Experience implementing SLOs, SLIs, error budgets, and reliability engineering best practices.
Strong troubleshooting and root cause analysis skills.
Experience developing runbooks, operational documentation, and automated remediation workflows.
Excellent communication and collaboration skills.
Preferred Qualifications:
Experience with Moogsoft or similar AIOps/observability platforms.
Experience implementing self-healing infrastructure and intelligent automation.
Familiarity with cloud platforms, containerization, and orchestration technologies.
Experience with AI-assisted operations (AIOps), monitoring, and predictive incident management.
If you are a Site Reliability Engineer with expertise in Java, Kafka, DevOps, and production automation, we''d love to hear from you. Apply today with your updated resume.