SRE

Woonsocket, RI, US • Posted 13 hours ago • Updated 13 hours ago
Contract W2
1 Year
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Site Reliability Engineer
  • Incident Management
  • P2 Incident
  • Kubernetes
  • Python
  • Java
  • GCP
  • P1 Incident

Summary

Role:- SRE
Remote - Woonsocket, RI hybrid

Job Description/ Responsibilities
8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility
Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents structured leadership updates, not just participant involvement
Experience tuning and validating time-series anomaly detection models in a production observability context this is a Required qualification, not a preferred one; anomaly-based detection is a core function of this role
Strong programming proficiency in Python, React, and Java at production quality capable of writing operational tooling that other engineers will rely on
Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services
Deep observability platform experience: Prometheus, Grafana, OpenTelemetry, and at least two of the log aggregation solution (Loki, Splunk, Elasticsearch)
Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments
Strong cloud platform expertise in Google Cloud Platform (Google Cloud Platform) and Rancher K3s.
Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI-assisted tooling and development.
Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipeline : Apache Airflow and Tidal.
Preferred Qualifications
Experience owning Production Readiness Reviews or service launch gates.
Strong proficiency in transforming large-scale operational and telemetry data into actionable business insights using SQL-based analytics, and reporting frameworks: Google BigQuery, PostgreSQL.
Hands-on chaos or fault injection experience.
TIC (Technical Incident Commander) certification or equivalent structured incident command training
Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact
LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) design or implementation experience
Experience with streaming data platforms: Kafka.
Experience with service mesh and traffic management: Istio, Envoy.
Infrastructure-as-code proficiency at production scale: Terraform or Ansible

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10217521
  • Position Id: 9083546
  • Posted 13 hours ago
Contact the job poster
Johnson Jose

Johnson Jose

Technical Recruiter @ Ztek Consulting
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Woonsocket, Rhode Island

Today

Full-time

USD 92,700.00 - 203,940.00 per year

Somerville, Massachusetts

Today

Full-time

USD 160,000.00 - 200,000.00 per year

Holyoke, Massachusetts

Today

Full-time

USD 134,000.00 - 170,000.00 per year

Remote

4d ago

Easy Apply

Contract, Third Party

$63 - $65

Search all similar jobs