Senior Site Reliability Engineer (SRE)

Remote in Springville, UT, US • Posted 5 hours ago • Updated 5 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

🛠️ Calibrating flux capacitors...

Job Details

Skills

  • Amazon EKS
  • Computer Networking
  • Load Balancing
  • Routing
  • Service Level
  • Software Engineering
  • Access Control
  • DevOps
  • Systems Engineering
  • Identity Management
  • Encryption
  • Management
  • Amazon Web Services
  • Virtual Private Cloud
  • Amazon CloudFront
  • Amazon Route 53
  • Amazon RDS
  • Remote Desktop Services
  • Amazon S3
  • Terraform
  • Grafana
  • PostgreSQL
  • Database
  • SaaS
  • Customer Facing
  • Disaster Recovery
  • Documentation
  • IaaS
  • Scalability
  • Performance Tuning
  • Kubernetes
  • Application Support
  • Cloud Security
  • Operational Risk
  • Root Cause Analysis
  • Technical Writing
  • Process Improvement
  • Collaboration
  • Cloud Computing
  • Production Support
  • Incident Management

Summary

Description

The Senior Site Reliability Engineer (SRE) will implement, secure, and operate the cloud infrastructure that supports CenCore Group's proprietary enterprise SaaS platform. This role is responsible for maintaining a scalable, highly available, secure, and reliable cloud environment as the platform grows and supports enterprise customers.

Key Responsibilities
  • Manage, maintain, and improve AWS-based cloud infrastructure supporting enterprise SaaS operations.
  • Operate and support Kubernetes environments, including Amazon EKS.
  • Own platform reliability, scalability, availability, disaster recovery readiness, and operational resilience.
  • Design and support cloud networking, load balancing, routing, traffic management, and related infrastructure components.
  • Implement and maintain monitoring, alerting, logging, and observability solutions to support proactive issue detection and response.
  • Establish and document operational standards, Service Level Objectives (SLOs), incident response processes, and reliability best practices.
  • Partner with software engineering and product teams to improve application performance, platform stability, and deployment reliability.
  • Apply security best practices across IAM, secrets management, encryption, vulnerability remediation, access controls, and production operations.
  • Support production operations, troubleshoot critical issues, and participate in incident resolution as needed.

Requirements

Required Qualifications
  • Professional experience supporting cloud infrastructure, site reliability, DevOps, platform engineering, or systems engineering functions.
  • Hands-on experience with AWS cloud services and production cloud operations.
  • Experience administering or operating Kubernetes environments.
  • Working knowledge of infrastructure reliability, availability, scalability, incident response, and operational support practices.
  • Experience implementing monitoring, logging, alerting, or observability tools.
  • Ability to troubleshoot complex production issues and coordinate resolution across technical teams.
  • Strong understanding of cloud security fundamentals, including identity and access management, encryption, secrets management, and vulnerability remediation.
  • Ability to document technical processes, standards, and operational procedures.

Preferred Qualifications
  • Experience with AWS services such as EKS, ALB, VPC, CloudFront, Route 53, RDS/Aurora, S3, and IAM.
  • Experience with Terraform or other Infrastructure as Code tools.
  • Experience with monitoring platforms such as Datadog, CloudWatch, Grafana, Prometheus, or similar tools.
  • PostgreSQL administration, performance tuning, or database operations experience.
  • Experience supporting enterprise SaaS, cloud-native applications, or customer-facing production platforms.
  • Experience developing disaster recovery, operational readiness, or production support documentation.

Skills / Competencies
  • Cloud infrastructure operations and automation
  • Platform reliability, scalability, and performance optimization
  • Kubernetes administration and containerized application support
  • Monitoring, observability, and incident response
  • Cloud security and operational risk awareness
  • Technical troubleshooting and root cause analysis
  • Cross-functional collaboration with engineering, product, and operations teams
  • Clear technical documentation and process improvement

Work Environment and Physical Requirements
This role is primarily performed in a professional office or remote technology environment, depending on business needs and position requirements. Work involves regular use of a computer, collaboration tools, and cloud-based systems. The position may require participation in production support, incident response, or after-hours troubleshooting as needed. Physical requirements are generally sedentary and include prolonged periods of sitting, computer use, and communicating with internal teams.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 80184262
  • Position Id: 3a705307997a39713145b908bd925f68
  • Posted 5 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Springville, Utah

Today

Full-time

Remote or Lehi, Utah

Today

Full-time

USD 160,000.00 - 180,000.00 per year

Pennsylvania

Today

Full-time

Remote

Today

Full-time

USD 110,000.00 - 140,000.00 per year

Search all similar jobs