Senior Site Reliability Engineer

Menlo Park, CA, US • Posted 9 hours ago • Updated 11 minutes ago
Full Time
On-site
GBP - GBP
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • IaaS
  • Operational Excellence
  • Scalability
  • Service Delivery
  • Incident Management
  • Service Level
  • Hardening
  • Capacity Management
  • Disaster Recovery
  • Root Cause Analysis
  • Recovery
  • Workflow
  • ARM
  • Mentorship
  • Collaboration
  • Accountability
  • Continuous Improvement
  • DevOps
  • Amazon Web Services
  • Google Cloud Platform
  • Google Cloud
  • Cloud Computing
  • Orchestration
  • Kubernetes
  • Continuous Integration
  • Continuous Delivery
  • Scripting
  • Python
  • Bash
  • Terraform
  • Grafana
  • Problem Solving
  • Conflict Resolution
  • Management
  • Microsoft Azure
  • Regulatory Compliance
  • FedRAMP
  • NIST 800-53
  • Sales

Summary

We are partnering with a forward-thinking organisation within the technology industry, renowned for its innovative approach to cloud infrastructure and operational excellence. Our Client fosters a dynamic and inclusive culture, dedicated to growth, continuous improvement, and delivering high-quality solutions to its clients. Committed to attracting talented professionals, they offer a stimulating environment where expertise is valued, and career development is actively supported.

Role Overview

An exciting opportunity has arisen to join Our Client as a Senior Site Reliability Engineer. This strategic role is vital as the organisation continues to expand its multi-cloud operations across AWS, Azure, and Google Cloud Platform platforms. The successful candidate will be instrumental in driving the reliability, scalability, and performance of critical systems, ensuring seamless service delivery and operational resilience. This position offers a unique chance to influence architectural decisions, pioneer automation practices, and lead incident management efforts that directly impact organisational success.

Key Responsibilities


  • Develop and oversee the infrastructure reliability strategy across cloud environments
  • Enhance observability through logging, monitoring, and alerting systems
  • Establish and manage Service Level Objectives (SLOs) and Service Level Agreements (SLAs) for key services
  • Lead performance optimisation, system hardening, capacity planning, and disaster recovery planning
  • Manage the incident lifecycle from initial detection to post-incident review and root cause analysis
  • Automate deployment, scaling, and recovery workflows to improve efficiency
  • Contribute to infrastructure as code initiatives using tools such as Terraform, CloudFormation, or ARM templates
  • Mentor and guide junior engineers and collaborate with cross-functional teams to foster best practices
  • Promote a culture of ownership, accountability, and continuous improvement

Essential Skills & Experience


  • Minimum of 5 years' experience in SRE, DevOps, or infrastructure engineering
  • Proven experience with large-scale systems in multi-cloud environments, particularly AWS and Google Cloud Platform
  • Strong understanding of cloud-native architecture, container orchestration with Kubernetes, and CI/CD pipelines
  • Proficiency in scripting languages such as Python or Bash
  • Hands-on experience with infrastructure automation tools (e.g., Terraform)
  • Familiar with monitoring and observability platforms like Prometheus, Grafana, Datadog, or ELK
  • Excellent problem-solving skills and the ability to work effectively under pressure
  • Clear communicator capable of translating technical information to diverse audiences

Desirable Skills & Experience


  • Experience managing infrastructure within Azure
  • Hands-on knowledge of Google SecOps
  • Understanding of compliance frameworks such as FedRAMP or NIST 800-53
  • Experience engaging with customers and prospects during technical integrations or pre-sales discussions

Next Steps

If you meet these criteria and are eager to make a meaningful impact within a reputable organisation that values innovation and professional development, we would love to hear from you. Please submit your CV to be considered for this exciting opportunity.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91114871
  • Position Id: 180838
  • Posted 9 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

San Mateo, California

Today

Full-time

USD 220,000.00 - 339,974.00 per year

Cupertino, California

Today

Full-time

San Mateo, California

Today

Full-time

USD 130,000.00 - 170,000.00 per year

San Jose, California

Today

Full-time

USD 87,600.00 - 186,000.00 per year

Search all similar jobs