Manager of DevOps Software Engineering

Hybrid in Chicago, IL, US • Posted 7 days ago • Updated 7 days ago
Full Time
Hybrid
$180,000 - $200,000/yr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • DevOps
  • Production Support
  • SLA
  • Kafka
  • Containers
  • Middleware
  • Monitoring
  • Observability

Summary

***We are unable to sponsor as this is a permanent full-time role***

***Hybrid, 3 days onsite, 2 days remote***

Responsibilities:

  • Lead L1 and L2 support engineers in all incident response activities including triage, investigation, coordination, resolution, closure, and post-incident reporting.
  • Oversee technical analysis of environment incidents across application deployments, middleware, and platform layers while coordinating response activities with internal engineering, platform, and application development teams.
  • Serve as Tier 3 escalation point for complex incidents beyond L2 capability triaging, directing, and driving resolution across Platform (k8s, Kafka, TFE), S&I (deployment, middleware, storage, network), Security (Vault, certs, secrets), and App Dev teams.
  • Drive post-incident reviews for all P1 and P2 incidents, ensuring root cause is identified, documented, and actioned not filed.
  • Define, publish, and enforce SLA targets across all severity levels
  • Monitor SLA compliance in real time; escalate breaches immediately and report trends to leadership on a sprint cadence.
  • Publish monthly SLA compliance reports to leadership with trend analysis and improvement actions.
  • Lead alert tuning and noise reduction initiatives across monitoring toolsets on-call engineers are paged for situations requiring human judgement, not system noise.
  • Track and publish the alert-to-incident ratio each sprint; hold the team accountable to a visible and improving trend.
  • Lead automation and tooling initiatives to reduce toil, accelerate triage, and eliminate manual steps from the support workflow.
  • Ensure accurate, complete documentation for every incident symptoms, steps taken, diagnostics, resolution, and RCA where applicable.
  • Own the runbook library every novel resolution produces a runbook published to L1 before the incident is closed; coverage gaps are tracked and closed sprint-on-sprint.
  • Lead a team of 6 10 L1 and L2 support engineers and contingent labor within the Environment Operations function.
  • Manage team scheduling to ensure full coverage of production support windows including on-call rotations, shift handoffs, and escalation availability for 247 support responsibilities.

Qualifications:

  • Minimum 5 years of hands-on environment operations, production support, or infrastructure operations experience, including interdisciplinary experience across four or more of the following: application deployment pipelines, container platform operations, middleware support, incident management, monitoring and observability, configuration management, release engineering, platform operations, or scripting and automation.
  • Technical experience and comprehensive knowledge of production environment failure modes including deployment failures, configuration drift, platform instability, and integration breakdowns and the methodologies used to diagnose and resolve them.
  • Demonstrated experience defining and enforcing SLA frameworks in a tiered support model (L1/L2/L3 or equivalent).
  • Shift work and on-call availability required including 247 on-call response capacity and availability during planned and emergency maintenance windows.
  • Previous people management or team lead experience required; formal people management experience strongly preferred.
  • Deployment & Pipeline tooling: Harness (continuous delivery pipelines, deployment verification, rollback automation), Jenkins (CI/CD pipeline management, job configuration, build troubleshooting), GitHub (branching strategies, pull request workflows, pipeline integration).
  • Container & orchestration platforms: Kubernetes (k8s) pod lifecycle management, namespace operations, log retrieval, resource troubleshooting, and coordination with Platform teams on cluster-level issues.
  • Messaging & streaming platforms: Apache Kafka topic management, consumer group monitoring, lag analysis, and escalation to Platform for broker-level issues.
  • Secrets & configuration management: HashiCorp Vault secrets retrieval, token/lease troubleshooting, policy review, and escalation to Security teams for certificate and secrets rotation.
  • Monitoring & observability: Proficiency in at least two production monitoring toolsets (e.g. Splunk, Dynatrace, Datadog, AppDynamics, PrometheGrafana) alert triage, dashboard interpretation, log analysis, and tuning requests.
  • Middleware platforms: Working knowledge of middleware infrastructure including application servers, messaging brokers, storage integrations, and network-layer dependencies sufficient to triage, gather diagnostics, and route correctly to L3.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: napil006
  • Position Id: 9066824
  • Posted 7 days ago
Contact the job poster
DG

Dillon Grooss

Recruiter @ Request Technology, LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Chicago, Illinois

13d ago

Easy Apply

Full-time

$170,000 - $200,000

Chicago, Illinois

Today

Full-time

USD 160,000.00 - 200,000.00 per year

Chicago, Illinois

Today

Full-time

USD 160,000.00 - 200,000.00 per year

Chicago, Illinois

Today

Full-time

USD 160,000.00 - 210,000.00 per year

Search all similar jobs