Only W2: SRE / Production Reliability Engineer, Woonsocket, RI


Tror
Dice Job Match Score™
🔢 Crunching numbers...
Job Details
Skills
- SRE/DevOps
- Incident Management
- Observability
- Kubernetes
- GCP
- and production monitoring
Summary
Experience: 8+ Years
Location: Woonsocket, RI
-
Own reliability and performance of critical production applications.
-
Act as an Incident Commander (IC) during P1/P2 production incidents.
-
Lead incident response, root-cause analysis, postmortems, and reliability improvements.
-
Define and manage SLIs, SLOs, and error budgets.
-
Build and improve monitoring, alerting, and observability solutions.
-
Tune and validate time-series anomaly detection models for production monitoring.
-
Develop automation and operational tools using Python, Java, and React.
-
Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
-
Work with engineering and operations teams to improve system reliability and reduce manual work.
-
Support large-scale deployments and manage production risks such as configuration drift and blast radius.
-
8+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
-
Hands-on experience as an Incident Commander for P1/P2 incidents.
-
Strong experience with time-series anomaly detection models in production observability — mandatory.
-
Strong production-level programming skills in:
-
Python
-
Java
-
React
-
-
Strong experience with SLI, SLO, and error budgets.
-
Hands-on observability experience with:
-
Prometheus
-
Grafana
-
OpenTelemetry
-
At least 2 log platforms such as Loki, Splunk, or Elasticsearch
-
-
Strong Google Cloud Platform experience.
-
Strong Kubernetes operational experience.
-
Experience with Rancher K3s.
-
Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
-
Experience with production-scale distributed systems and on-call support.
-
Production Readiness Reviews / service launch experience.
-
BigQuery and PostgreSQL.
-
Chaos Engineering / fault injection.
-
TIC/Technical Incident Commander certification.
-
Healthcare, pharmacy, retail, or other high-availability environments.
-
LLM/GenAI for incident management, alert summarization, or runbook recommendations.
-
Kafka.
-
Istio / Envoy.
-
Terraform / Ansible.
- Dice Id: 91135853
- Position Id: 673-38728-1787769724
- Posted 4 hours ago
Company Info
TROR is an artificial intelligence consultancy specializing in developing powerful and customized Al solutions for business. With top Al Experts we take pride in providing the best cutting-edge Al consultancy. Our years of experience in various industries helps us to develop and implement bespoke Al solutions for businesses. Our on demand Al products have helped over 100 companies drive transformational results.
The solutions we bring on your table meet the highest industry standards and quality, effectively and efficiently resolving your issues and optimizing the way you want to move forward in the market. Through our customer centric approach, we ensure that we are always there for our valuable customers by offering them satisfactory solutions for guaranteed results.


Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs