Site Reliability Engineer

Hybrid in Atlanta, GA, US • Posted 1 day ago • Updated 1 day ago
Full Time
No Travel Required
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • SRE
  • reliability
  • automation
  • scripting

Summary

Role: Site Reliability Engineer

Work location: Hybrid, onsite at least 3 days per week in Midtown, Atlanta, GA.

Duration: Full time (direct hire)

 

Role Summary
Lead the reliability, scalability, security, and operational excellence of customer-facing platforms across Azure, Google Cloud Platform, and Kubernetes environments. Drive production stability through automation, observability, incident management, and continuous improvement initiatives.

 

Key Responsibilities

 

  • Lead platform reliability, availability, and performance initiatives.
  • Design and support cloud infrastructure in Azure and Google Cloud Platform.
  • Manage and optimize Kubernetes environments and containerized applications.
  • Implement observability and monitoring using Splunk, AppDynamics, and cloud-native tools.
  • Support Cloudflare, Zscaler, SQL Server, RabbitMQ, and enterprise networking components.
  • Lead major incident response, RCA, and problem management activities.
  • Develop automation and self-healing solutions to improve operational efficiency.
  • Collaborate with Engineering, Product, Security, and Infrastructure teams to enhance customer experience and platform stability.
  • Serve as a technical escalation point for critical production and customer issues.

 

Required Skills

 

  • 5+ years of experience in SRE, DevOps, Cloud Operations, or Infrastructure Engineering.
  • Strong expertise in Azure, Google Cloud Platform, Kubernetes, Cloudflare, Splunk, AppDynamics, SQL Server, RabbitMQ, and Zscaler.
  • Solid networking knowledge (DNS, TCP/IP, HTTP/S, CDN, WAF, Load Balancing, SSL/TLS, Firewalls, VPNs).
  • Experience with automation and scripting (Python, PowerShell, Bash, Terraform).
  • Strong customer-facing communication and stakeholder management skills.

 

Core Principles

 

  • Automation First – Eliminate manual effort through automation and self-healing systems.
  • Observability-Driven Operations – Leverage logs, metrics, traces, and analytics to proactively identify and resolve issues.
  • AI-Powered Reliability – Utilize AI and operational intelligence to accelerate detection, diagnosis, and remediation.
  • Customer-Centric Mindset – Prioritize customer experience, stability, and business outcomes.
  • Operational Excellence – Continuously improve reliability, scalability, and security.

 

Success Measures

 

  • Service Availability & Uptime
  • SLA/SLO Compliance
  • MTTR Reduction
  • Incident Reduction
  • Automation Adoption
  • Customer Satisfaction (CSAT)
  • Platform Performance & Stability Improvements

 

Best Regards,

-------

David Roy | #LI-DR1 Accounts Manager – US Staffing | Charter Global Inc. |    

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: chartpro
  • Position Id: 30575-13826-1790626699
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Atlanta, Georgia

•

Yesterday

Easy Apply

Full-time

90,000 - 120000

Atlanta, Georgia

•

Today

Full-time

USD 200,000.00 - 225,000.00 per year

Atlanta, Georgia

•

6d ago

Full-time

USD 211,000.00 - 317,000.00 per year

Hybrid in Alpharetta, Georgia

•

Today

Easy Apply

Third Party, Contract

Depends on Experience

Search all similar jobs