Site Reliability Engineer (SRE)

Remote • Posted 8 hours ago • Updated 8 hours ago
Full Time
No Travel Required
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Amazon Web Services
  • Bash
  • HIPAA
  • Health Care
  • Jenkins
  • Kubernetes
  • Microsoft Azure
  • NoSQL
  • Production Engineering
  • Python
  • RBAC
  • Regulatory Compliance
  • Root Cause Analysis
  • SQL
  • Software Development Methodology
  • Terraform
  • Windows PowerShell
  • Workflow
  • Software Engineering
  • Scripting
  • CircleCI
  • Grafana
  • IaaS
  • DevOps
  • Database
  • GitHub
  • Datadog
  • Prometheus
  • CloudWatch

Summary

Site Reliability Engineer (SRE)  

We are seeking a highly technical Site Reliability Engineer to build, operate, automate, and scale cloud-native production platforms supporting mission-critical applications in regulated healthcare environments. 

Core Responsibilities 

  • Own production reliability, availability, scalability and performance across cloud infrastructure and application services. 

  • Define and implement SLOs, SLIs, error budgets, alerting strategies and incident response processes. 

  • Partner with software engineering teams to embed reliability, resiliency and operational readiness into the SDLC. 

  • Lead root cause analysis, incident management and post-mortem activities. 

  • Automate infrastructure provisioning, deployments and operational workflows. 

Required Technical Skills 

  • Cloud Platforms: AWS, Azure or Google Cloud Platform. 

  • Containers & Orchestration: Docker, Kubernetes, Helm. 

  • Infrastructure as Code: Terraform or CloudFormation. 

  • CI/CD: Jenkins, GitHub Actions, GitLab CI/CD or CircleCI. 

  • Databases: Operational support, backup, recovery and performance tuning of SQL/NoSQL databases. 

  • Scripting: Python, Bash or PowerShell. 

Observability & Reliability Engineering 

  • Hands-on experience with Datadog, Prometheus, Grafana, CloudWatch, Splunk or Dynatrace. 

  • Build dashboards, logging pipelines, tracing solutions and proactive monitoring frameworks. 

  • Experience diagnosing latency, throughput, capacity and availability issues in distributed systems. 

Security & Compliance 

  • Knowledge of HIPAA or equivalent healthcare regulations. 

  • Secrets management using Vault, AWS Secrets Manager or Azure Key Vault. 

  • Strong understanding of IAM, RBAC, network security and least-privilege access models. 

Qualifications 

  • 3-6+ years of experience in SRE, DevOps or Production Engineering. 

  • Experience supporting highly available customer-facing production services. 

  • Strong troubleshooting and incident management capabilities. 

 

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91126544
  • Position Id: 9094906
  • Posted 8 hours ago
Contact the job poster
NK

Nitesh Kumar

Recruiter @ StatusNeo Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote or Atlanta, Georgia

•

Today

Full-time

USD 120,000.00 - 175,000.00 per year

Remote

•

30+d ago

Easy Apply

Full-time

$180,000 - $215,000

Remote

•

Today

Full-time

USD 130,000.00 - 160,000.00 per year

Remote

•

Today

Full-time

Search all similar jobs