Site Reliability Engineer (SRE) Production Services

Pittsburgh, PA, US • Posted 6 hours ago • Updated 6 hours ago
Contract W2
12 Months
No Travel Required
On-site
$55 - $58/hr
Company Branding Image
Fitment

Dice Job Match Score™

👤 Reviewing your profile...

Job Details

Skills

  • Site Reliability Engineering
  • SRE
  • Java
  • Spring Boot
  • Apache Kafka
  • DevOps
  • CI/CD
  • CI/CD Automation
  • Production Support
  • Production Services
  • Automation
  • Reliability Engineering
  • Service Level Objectives
  • SLO
  • SLI
  • Error Budgets
  • Observability
  • Monitoring
  • Dashboards
  • Incident Management
  • Problem Management
  • Runbooks
  • Self-Healing
  • AIOps
  • Moogsoft
  • Root Cause Analysis
  • Operational Automation
  • Event-Driven Architecture
  • Batch Processing
  • Intelligent Operations
  • Platform Engineering

Summary

Job Title: Site Reliability Engineer (SRE) – Production Services

Location: Pittsburgh, PA (Onsite)

Client Address: 500 Grant Street, Pittsburgh, PA 15219

Employment Type: Contract

Rate: $55–58/hr

Experience Required: 8–10 Years

Note: Local candidates only.

We are seeking an experienced Site Reliability Engineer (SRE) to join our Production Services team. This role is ideal for professionals with strong expertise in Java Spring Boot, Apache Kafka, DevOps, and CI/CD automation who are passionate about building resilient, highly available, and self-healing production platforms.

The ideal candidate will drive automation, improve operational efficiency, implement observability solutions, and enhance production reliability through modern SRE and DevOps practices.

Top Required Skills:

  • Java Spring Boot

  • Apache Kafka

  • DevOps & CI/CD Automation

Key Responsibilities:

  • Automate high-volume production support requests and operational workflows to improve efficiency and reduce manual effort.

  • Develop self-service capabilities and automated remediation solutions for recurring operational tasks.

  • Design resilient operational workflows with auditability, consistency, and fault tolerance.

  • Implement automated retry mechanisms and intelligent backoff strategies for recurring production failures.

  • Define, implement, and manage Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical applications and batch processes.

  • Apply error budget principles to support release management and reliability improvements.

  • Improve batch processing reliability through standardized recovery patterns and proactive monitoring.

  • Build observability dashboards to monitor incidents, failure rates, repeat issues, automation coverage, and production health.

  • Enhance operational reporting across incidents, changes, and problem management processes.

  • Develop comprehensive runbooks and convert manual operational procedures into automated workflows.

  • Drive self-service capabilities for common operational requests and recurring production activities.

  • Implement self-healing capabilities to automatically detect and remediate production issues.

  • Optimize monitoring and alerting platforms, including Moogsoft, to reduce alert fatigue and improve signal quality.

  • Leverage automation and AI-driven operational practices to proactively resolve recurring production issues.

  • Collaborate with Development, Infrastructure, DevOps, and Operations teams to continuously improve platform reliability.

Required Qualifications:

  • 8–10 years of experience in Site Reliability Engineering, Production Support, DevOps, or Platform Engineering.

  • Strong hands-on experience with Java and Spring Boot.

  • Strong expertise in Apache Kafka and event-driven architectures.

  • Experience with DevOps practices and CI/CD automation.

  • Experience building automation solutions using scripting and infrastructure automation tools.

  • Strong knowledge of production monitoring, observability, logging, and incident management.

  • Experience implementing SLOs, SLIs, error budgets, and reliability engineering best practices.

  • Strong troubleshooting and root cause analysis skills.

  • Experience developing runbooks, operational documentation, and automated remediation workflows.

  • Excellent communication and collaboration skills.

Preferred Qualifications:

  • Experience with Moogsoft or similar AIOps/observability platforms.

  • Experience implementing self-healing infrastructure and intelligent automation.

  • Familiarity with cloud platforms, containerization, and orchestration technologies.

  • Experience with AI-assisted operations (AIOps), monitoring, and predictive incident management.

If you are a Site Reliability Engineer with expertise in Java, Kafka, DevOps, and production automation, we''d love to hear from you. Apply today with your updated resume.

 

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10411276
  • Position Id: 9031595
  • Posted 6 hours ago

Company Info

About K-Tek Resourcing LLC

Vision

To be a trusted partner and advisor to our customers

Mission

At K-Tek we believe in understanding the specific needs of the customer and tailor-creating innovative solutions to meet these needs. We invest in our employees and customers. We build a relation of trust with our customers through empathy, solutions and being the first time right.

Who We Are & What We Do

K-Tek Resourcing is a consulting organization with offices in Houston TX and St. Paul, MN. It is supported by 2 global delivery centers, located in India. With its global employee strength of over 250, K-Tek has been supporting its clients for over 9 years. We have been consistently achieving a growth of 30% Year on Year. We have an extensive experience of working in domains including BFSI, Retail, Healthcare and Pharma, Oil & Gas, Travel & Hospitality and Insurance. The technologies we service are IT Infrastructure, Mobile Technologies, Cloud & Big Data Solutions. We understand the needs of our customers and provide them with customized solutions and resources with the tenet of being the "First Time Right".

Values

-Commitment to our customers success through Integrity

-Excellence through Quality

-Growth through customer value creation

About_Company_One
Contact the job poster
DG

Darpan Gome

Recruiter @ K-Tek Resourcing LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Atlanta, Georgia

Today

Easy Apply

Contract

60 - 65

Hybrid in Dallas, Texas

Today

Easy Apply

Contract

48 - 50

Search all similar jobs