Site Reliability Engineer

Hybrid in Chicago, IL, US • Posted 8 days ago • Updated 8 days ago
Full Time
No Travel Required
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Site Reliability Engineer
  • Python
  • GCP
  • Platform Engineering
  • Internal Development Platform
  • Kubernetes
  • Artificial Intelligence
  • Terraform

Summary

We are seeking a Staff Site Reliability Engineer to serve as the foundational Technical Lead for our Platform Engineering SRE organization. In this role, you will be the primary architect and visionary for the core technology foundations that underpin the world's leading derivatives marketplace.
As the technical lead for all SRE sub-teams, you will bridge the gap between high-level business goals and deep technical implementation, ensuring our Google Cloud Platform-native stack provides the mission-critical intersection of ultra-low latency and absolute reliability required for high-volume financial ecosystems.
As the Staff SRE, your goal is to evolve our platform from "infrastructure as a service " to "reliability as a product. " You will be responsible for the technical roadmap of our entire SRE domain, mentoring senior engineers and setting the global standard for operational excellence across our Python, Kafka, and Kubernetes shop.

What You Will Do

  • Technical Vision & Roadmap: Define the 12–18 month technical strategy for the Platform SRE teams, focusing on the evolution of our global footprint and self-service capabilities.
  • Architectural Authority: Act as the final technical authority for major infrastructure changes involving Google Cloud Platform, GKE, and our mission-critical Kafka event bus.
  • Incident Command & Systemic Resilience: Lead the response for complex, cross-functional outages and drive a "blameless " culture that prioritizes systemic, code-based fixes over manual intervention.
  • Internal Development Platform (IDP): Architect and oversee the building of high-level abstractions in Python to mask underlying complexity, providing a seamless "Golden Path " for our global technology stack.
  • Reliability Governance: Standardize SLIs, SLOs, and Error Budgets across all platform teams, ensuring they are technically rigorous and directly tied to market integrity.
  • Engineering Mentorship: Level up the entire SRE organization through design reviews, architectural "office hours, " and fostering an environment of continuous technical evolution.


What We're Looking For

  • Strategic AI Integration: A mastery of leveraging Generative AI and Agentic workflows (e.g., Gemini) to build self-healing infrastructure and sophisticated automated troubleshooting frameworks.
  • Software Engineering Mastery: Expert-level proficiency in Python (and ideally Go) to build production-grade distributed systems and custom Kubernetes operators.
  • Cloud-Native Leadership: Deep-seated expertise in Google Cloud Platform (Networking, IAM, GKE) and the ability to scale Kafka clusters for high-throughput, low-latency financial environments.
  • Advanced IaC & GitOps: Mastery of Terraform module design and ArgoCD for managing immutable infrastructure at an enterprise scale.
  • Distributed Systems Theory: A rigorous understanding of non-linear system behaviors, distributed consensus, and the nuances of high-concurrency architectures.
  • Executive Communication: The ability to translate sophisticated technical debt and architectural risks into clear business outcomes for senior leadership.


Experience:

  • 10+ years in SRE, Systems Engineering, or Software Engineering roles within high-pressure environments.
  • 3+ years in a Staff, Principal, or Tech Lead capacity overseeing multiple teams or complex platform domains.
  • Proven Track Record: Experience leading large-scale cloud migrations or re-architecting core messaging/compute platforms in a regulated environment.
  • Certifications: Google Cloud Platform Professional Cloud Architect or Kubernetes (CKA/CKAD).
  • Full-Stack Exposure: Proficiency in Node.js or modern front-end frameworks.
  • Domain Expertise: Experience in Financial Markets or highly regulated, high-concurrency environments.
  • Agile Integration: Comfort working within Agile frameworks and highly collaborative software development lifecycles.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10174401
  • Position Id: 9076089
  • Posted 8 days ago
Contact the job poster
Azaad Babu

Azaad Babu

Recruiter @ Informatic Technologies
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Chicago, Illinois

11d ago

Full-time

Chicago, Illinois

Today

Full-time

USD 160,000.00 - 210,000.00 per year

Oak Brook, Illinois

Today

Full-time

USD 129,700.00 - 226,900.00 per year

Remote

Today

Full-time

Search all similar jobs