Site Reliability Engineer - (LATAM - Columbia, Mexico , Costa Rica)

Remote • Posted 3 hours ago • Updated 3 hours ago
Contract Independent
Contract W2
12 Months
No Travel Required
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • AWS
  • Amazon Web Services
  • Google Cloud Platform
  • Kubernetes
  • Reliability Engineering
  • Performance Tuning
  • Python
  • Splunk
  • Migration

Summary

Senior Site Reliability Engineer (SRE)

Remote from ONLY LATAM  - (Colombia, Mexico, Costa Rica, Brazil, Peru)

We are looking for a Senior Site Reliability Engineer to build, operate, and improve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. This is a hands-on technical leadership role focused on AWS/Google Cloud Platform, Kubernetes, observability, automation, and reliability engineering.

Key Responsibilities

  • Design and implement reliability strategies for distributed systems across AWS and Google Cloud Platform.
  • Define and monitor SLIs, SLOs, error budgets, and reliability metrics.
  • Build and enhance monitoring, logging, tracing, alerting, and observability solutions.
  • Lead incident response, root cause analysis, and postmortems.
  • Improve system performance, scalability, resiliency, and operational readiness.
  • Automate operational processes and reduce manual toil.
  • Guide engineering teams on reliability architecture, capacity planning, and non-functional requirements.

Required Skills

  • 7+ years in SRE, Cloud Engineering, DevOps, or Platform Engineering.
  • Strong production experience with AWS and/or Google Cloud Platform.
  • Hands-on expertise with Kubernetes (EKS/GKE).
  • Strong understanding of SRE principles, SLIs, SLOs, error budgets, and incident management.
  • Experience with observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, or Splunk.
  • Strong Terraform/IaC experience.
  • Proficiency in Python, Bash, or similar scripting languages.
  • Strong understanding of cloud networking, distributed systems, security, and performance optimization.

Preferred

  • Large-scale cloud migration or modernization experience.
  • Chaos engineering/resilience testing experience.
  • Istio/service mesh knowledge.
  • AWS/Google Cloud Platform certifications.
  • Experience in Agile, DevOps, or DevSecOps environments.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90773816
  • Position Id: 9076170
  • Posted 3 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote or Atlanta, Georgia

Today

Contract

$35 - $44 hourly

Remote

Today

Full-time

USD 125,000.00 - 145,000.00 per year

Remote

Today

Full-time

Remote

Today

Full-time

USD 151,500.00 - 252,500.00 per year

Search all similar jobs