SRE Lead


Drunix Solution Inc
Dice Job Match Score™
🔢 Crunching numbers...
Job Details
Skills
- SRE
Summary
Hi,
Greetings of the day!!
Job Location: Plano, TX / Charlotte, NC (3 Days onsite, 2 Days remote) Hybrid role.
Job Type: 12+ Months contract
- Production engineering and reliability leadership for messaging platforms
- Platform security, resilience engineering, and vulnerability remediation
- Ownership of large-scale, distributed messaging runtimes
- Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput
- Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters
- SLIs / SLOs focused on message delivery, latency, and availability
- Incident management, escalation, and postmortem culture
- Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput
- Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery)
- Building resilient messaging platforms capable of supporting real-time, event-driven workloads
- Supporting global production messaging environments with on-call rotation and escalation ownership
- Partnering with engineering, application, and security teams tensure reliability, scalability, and secure message transport
- Strong experience in Site Reliability Engineering / Production Engineering
- IBM MQ (queue managers, clustering, channels, DLQ management)
- Kafka / Confluent platform (topics, brokers, partitions, consumer groups)
- Large-scale distributed messaging systems and runtime management
- System reliability, scalability, and high availability design
- Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering)
- Incident management, root cause analysis, and problem management
- Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms
- Event and anomaly detection in high-volume systems
- Strong scripting/automation skills:
- Shell, Python, PowerShell
- Experience managing Linux/Unix and Windows production environments
- Event-driven architecture and messaging-based integration patterns
- Messaging platform security (TLS, certificates, channel auth, encryption)
- Vulnerability remediation and risk mitigation in production systems
- Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures
- Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads
- Kubernetes / containerized messaging platforms
- Experience with:
- Kafka ecosystem components (Schema Registry, Connect, Streams)
- IBM MQ advanced features (Native HA, clustering)
- AI-driven operations (AIOps), anomaly detection, or automated remediation
- Large-scale messaging modernization or migration programs
- Messaging or middleware certifications (IBM MQ, Kafka, or equivalent)
- Experience in regulated environments (e.g., financial services)
- Dice Id: 91172814
- Position Id: 9067533
- Posted 4 hours ago
Company Info
Mission:
Our mission is to empower businesses with innovative and scalable IT solutions. We strive to deliver excellence in software development, IT infrastructure, staffing, and healthcare systems that drive performance and digital transformation.
Vision:
To become a global leader in IT services by offering tailored, future-ready solutions that enable our clients to excel in an ever-evolving digital landscape. We envision a world where technology accelerates growth across all sectors.
Values:
We value integrity, innovation, and client success. Our team is committed to continuous learning, transparent collaboration, and delivering measurable results through technology that truly matters.

Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs