Site Reliability Engineer

Hybrid in Charlotte, NC, US • Posted 4 hours ago • Updated 4 hours ago
Contract W2
18 Months
No Travel Required
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • SRE
  • Site reliability engineer
  • python
  • IBM MQ
  • PowerShell
  • Grafana
  • Production Engineering
  • Kafka
  • Messaging Services

Summary

Site Reliability Engineer (SRE) – Messaging Services

We are seeking an experienced SRE – Messaging Services to drive reliability, observability, security, and operational excellence across IBM MQ and Kafka/Confluent environments.

Key Responsibilities:

  • Lead reliability engineering for large-scale IBM MQ and Kafka platforms.
  • Drive patching, EOL remediation, vulnerability remediation, and platform stabilization.
  • Define and implement SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems.
  • Enhance monitoring and observability for message flows, queue depths, Kafka lag, throughput, and latency.
  • Develop proactive fault detection and automated remediation strategies.
  • Support high-availability, resilient, and scalable messaging platforms.
  • Troubleshoot production issues including message backlogs, latency spikes, and connection failures.
  • Partner with engineering, application, infrastructure, and security teams.
  • Support global production environments and participate in on-call rotations.

Required Skills:

  • Strong experience in SRE / Production Engineering.
  • Hands-on expertise with IBM MQ and Kafka/Confluent.
  • Strong knowledge of distributed messaging, high availability, scalability, and reliability patterns.
  • Experience with Dynatrace, Splunk, Prometheus, Grafana, or similar observability tools.
  • Strong scripting/automation skills using Python, Shell, or PowerShell.
  • Experience with Linux/Unix and Windows production environments.
  • Knowledge of messaging security, TLS, certificates, encryption, and vulnerability remediation.
  • Strong troubleshooting, incident management, and RCA skills.

Preferred:

  • Experience with Kubernetes/containerized messaging platforms.
  • Knowledge of Kafka Schema Registry, Connect, and Streams.
  • Experience with IBM MQ clustering or Native HA.
  • Exposure to AIOps, anomaly detection, or automated remediation.
  • Experience with messaging modernization/migration programs.
  • Financial services or other regulated-industry experience preferred.

Candidate Requirements:

  • Strong communication and leadership skills.
  • Ability to work in high-pressure production environments.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91132487
  • Position Id: 9065098
  • Posted 4 hours ago
Contact the job poster
GV

Gaurav Verma

Recruiter @ Techgroup America Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Charlotte, North Carolina

Today

Easy Apply

Contract

$51.25 - $56.16

Charlotte, North Carolina

Today

Easy Apply

Full-time

$53 - $57 per hour

No location provided

Today

Full-time

USD 180,000.00 - 240,000.00 per year

Montana

Today

Full-time

USD 189,000.00 - 232,000.00 per year

Search all similar jobs