W2 - Site Reliability Engineer (SRE) Lead – IBM MQ / Kafka | Hybrid | Plano, TX / Charlotte, NC

Hybrid in Plano, TX, US • Posted 3 hours ago • Updated 3 hours ago
Contract W2
12 Months
75% Travel Required
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • IBM WebSphere MQ
  • Microsoft Windows
  • Middleware
  • Migration
  • Operational Excellence
  • Incident Management
  • Leadership
  • Linux
  • Management
  • Failover
  • Financial Services
  • Grafana
  • High Availability
  • Kubernetes
  • Messaging
  • Production Engineering
  • Python
  • Root Cause Analysis
  • Real-time
  • Recovery
  • Reliability Engineering
  • Risk Management
  • Artificial Intelligence
  • Apache Kafka
  • Budget
  • Splunk
  • Streaming
  • Clustering
  • Dynatrace
  • Encryption
  • Scalability
  • Scripting
  • Shell
  • TLS
  • Unix

Summary

Job Title: Site Reliability Engineer (SRE)
Job Location: Plano, TX / Charlotte, NC (3 Days onsite, 2 Days remote) Hybrid role.
Job Type: 12+ Months contract
 
 
 
Job Descriptions:
We are seeking an experienced Site Reliability Engineer (SRE) Lead - Messaging Services tdrive platform reliability, observability, and operational excellence across IBM MQ and Kafka environments.
 
This role combines:
  • Production engineering and reliability leadership for messaging platforms
  • Platform security, resilience engineering, and vulnerability remediation
  • Ownership of large-scale, distributed messaging runtimes
Key responsibilities include:
  • Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput
  • Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters
Implementing SRE best practices:
  • SLIs / SLOs focused on message delivery, latency, and availability
  • Incident management, escalation, and postmortem culture
  • Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput
  • Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery)
  • Building resilient messaging platforms capable of supporting real-time, event-driven workloads
  • Supporting global production messaging environments with on-call rotation and escalation ownership
  • Partnering with engineering, application, and security teams tensure reliability, scalability, and secure message transport
  • Strong experience in Site Reliability Engineering / Production Engineering
Hands-on expertise with:
  • IBM MQ (queue managers, clustering, channels, DLQ management)
  • Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
  • Large-scale distributed messaging systems and runtime management
Deep understanding of:
  • System reliability, scalability, and high availability design
  • Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering)
  • Incident management, root cause analysis, and problem management
Experience with:
  • Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms
  • Event and anomaly detection in high-volume systems
  • Strong scripting/automation skills:
  • Shell, Python, PowerShell
  • Experience managing Linux/Unix and Windows production environments
Knowledge of:
  • Event-driven architecture and messaging-based integration patterns
Understanding of:
  • Messaging platform security (TLS, certificates, channel auth, encryption)
  • Vulnerability remediation and risk mitigation in production systems
  • Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures
  • Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads
Familiarity with:
  • Kubernetes / containerized messaging platforms
  • Experience with:
  • Kafka ecosystem components (Schema Registry, Connect, Streams)
  • IBM MQ advanced features (Native HA, clustering)
Exposure to:
  • AI-driven operations (AIOps), anomaly detection, or automated remediation
  • Large-scale messaging modernization or migration programs
  • Messaging or middleware certifications (IBM MQ, Kafka, or equivalent)
  • Experience in regulated environments (e.g., financial services)
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91126058
  • Position Id: 9066908
  • Posted 3 hours ago
Contact the job poster
Anusha Chenna

Anusha Chenna

Recruiter @ Prohires
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Plano, Texas

Today

Full-time

Plano, Texas

Today

Full-time

Plano, Texas

Today

Full-time

Hybrid in Plano, Texas

Yesterday

Easy Apply

Full-time

Up to $120,000

Search all similar jobs