Site Reliability Engineer

• Posted 5 days ago • Updated 5 days ago
Full Time
On-site
Fitment

Dice Job Match Score™

🛠️ Calibrating flux capacitors...

Job Details

Skills

  • Management
  • Incident Management
  • Root Cause Analysis
  • Systems Design
  • Capacity Management
  • Disaster Recovery
  • Operational Excellence
  • Computer Science
  • Reliability Engineering
  • Production Engineering
  • DevOps
  • Linux Administration
  • Kubernetes
  • Terraform
  • Cloud Computing
  • Amazon Web Services
  • Microsoft Azure
  • Google Cloud Platform
  • Google Cloud
  • Python
  • Bash
  • Scripting
  • Programming Languages
  • Continuous Integration
  • Continuous Delivery
  • Grafana
  • Splunk
  • Problem Solving
  • Conflict Resolution

Summary

Responsibilities

  • Design, build, and maintain highly available and scalable infrastructure
  • Automate operational processes to improve efficiency and reduce manual intervention
  • Manage and support Kubernetes-based containerized environments
  • Develop and maintain Infrastructure as Code using Terraform and related tools
  • Build and enhance CI/CD pipelines to streamline deployment processes
  • Monitor system health, performance, and reliability across production environments
  • Lead incident response efforts and drive root cause analysis for production issues
  • Partner with engineering teams to improve system design, resilience, and observability
  • Implement best practices around monitoring, alerting, capacity planning, and disaster recovery
  • Continuously identify opportunities to improve reliability, performance, and operational excellence


Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience)
  • Experience in Site Reliability Engineering, Platform Engineering, Production Engineering, DevOps, or Infrastructure Engineering
  • Strong Linux systems administration experience
  • Hands-on experience with Kubernetes and containerized environments
  • Experience with Terraform or other Infrastructure as Code tools
  • Strong knowledge of cloud platforms such as AWS, Azure, or Google Cloud Platform
  • Proficiency in Python, Go, Bash, or similar scripting/programming languages
  • Experience building and supporting CI/CD pipelines
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK
  • Strong troubleshooting and problem-solving skills in large-scale production environments
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90922487
  • Position Id: 24602088
  • Posted 5 days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Full-time

Remote

Today

Full-time

USD 87,400.00 - 123,400.00 per year

No location provided

Today

Full-time

USD 230,000.00 - 250,000.00 per year

No location provided

Today

Full-time

USD 81,100.00 - 187,000.00 per year

Search all similar jobs