W2 Candidates ONLY
Title: Site Reliability Engineer (SRE)
Location: Plano, TX / Charlotte, NC (3 Days onsite, 2 Days remote) Hybrid role.
Duration: 12+ Months contract
Interview Process: 2 virtual rounds; onsite interview may also be requested
Client Type: Direct Client (MSP/VMS)
Note:
Please submit only qualified candidates who closely match the job requirements.
Genuine LinkedIn is must.
Local candidates for onsite/hybrid roles.
BOA experience first, followed by other banking/financial clients, then strong enterprise clients.
Avoid candidates with 6+ month employment gaps.
Prefer candidates with stable, long-term employment/contract history.
Speed is critical due to the MSP/VMS process.
Please share qualified candidates for the SRE Lead position with strong hands-on experience in IBM MQ and Kafka/Confluent, large-scale messaging/production engineering, SRE practices, monitoring/observability, incident management, RCA, SLI/SLO, HA/resiliency, and Shell/Python/PowerShell automation. Linux/Windows experience is required, while Kubernetes and Banking/Financial Services experience is preferred.
Job Descriptions:
We are seeking an experienced Site Reliability Engineer (SRE) Lead - Messaging Services tdrive platform reliability, observability, and operational excellence across IBM MQ and Kafka environments.
This role combines:
- Production engineering and reliability leadership for messaging platforms
- Platform security, resilience engineering, and vulnerability remediation
- Ownership of large-scale, distributed messaging runtimes
Key responsibilities include:
- Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput
- Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters
Implementing SRE best practices:
- SLIs / SLOs focused on message delivery, latency, and availability
- Incident management, escalation, and postmortem culture
- Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput
- Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery)
- Building resilient messaging platforms capable of supporting real-time, event-driven workloads
- Supporting global production messaging environments with on-call rotation and escalation ownership
- Partnering with engineering, application, and security teams tensure reliability, scalability, and secure message transport
- Strong experience in Site Reliability Engineering / Production Engineering
Hands-on expertise with:
- IBM MQ (queue managers, clustering, channels, DLQ management)
- Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
- Large-scale distributed messaging systems and runtime management
Deep understanding of:
- System reliability, scalability, and high availability design
- Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering)
- Incident management, root cause analysis, and problem management
Experience with:
- Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms
- Event and anomaly detection in high-volume systems
- Strong scripting/automation skills:
- Shell, Python, PowerShell
- Experience managing Linux/Unix and Windows production environments
Knowledge of:
- Event-driven architecture and messaging-based integration patterns
Understanding of:
- Messaging platform security (TLS, certificates, channel auth, encryption)
- Vulnerability remediation and risk mitigation in production systems
- Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures
- Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads
Familiarity with:
- Kubernetes / containerized messaging platforms
- Experience with:
- Kafka ecosystem components (Schema Registry, Connect, Streams)
- IBM MQ advanced features (Native HA, clustering)
Exposure to:
- AI-driven operations (AIOps), anomaly detection, or automated remediation
- Large-scale messaging modernization or migration programs
- Messaging or middleware certifications (IBM MQ, Kafka, or equivalent)
- Experience in regulated environments (e.g., financial services)
Equal Opportunity Employer/Veterans/Disabled
Benefit offerings available for our associates include medical, dental, vision, life insurance, short-term disability, additional voluntary benefits, an EAP program, commuter benefits, and a 401K plan. Our benefit offerings provide employees with the flexibility to choose the type of coverage that meets their individual needs. In addition, our associates may be eligible for paid leave including Paid Sick Leave or any other paid leave required by Federal, State, or local law, as well as Holiday pay where applicable. Disclaimer: These benefit offerings do not apply to client-recruited jobs and jobs that are direct hires to a client.
The Company will consider qualified applicants with arrest and conviction records in accordance with federal, state, and local laws and/or security clearance requirements, including, as applicable:
- The California Fair Chance Act
- Los Angeles City Fair Chance Ordinance
- Los Angeles County Fair Chance Ordinance for Employers
- San Francisco Fair Chance Ordinance
Senior Talent Acquisition
E-
STELLENT IT A Nationally Recognized Minority Certified Enterprise