Hiring Senior SRE / Production Reliability Engineer, 100% Remote


Tror
Dice Job Match Score™
🔢 Crunching numbers...
Job Details
Skills
- Site Reliability Engineering
- Production Support
- Monitoring
- Observability
- Incident Management
- Root Cause Analysis
- RCA
- Python
- Automation
- DNS
- Networking
- Load Balancing
- WAF
- Cloud
- Azure
- GCP
- API
- Microservices
Summary
Role: Senior SRE / Production Reliability Engineer
Location: 100% Remote
Experience: 10+ Years
We need an experienced SRE who can keep production systems stable, find problems before users are affected, troubleshoot incidents, and automate repetitive checks.
The person will support Salesforce CRM (GPS) and YAVA platforms.
Looking for a strong SRE who can monitor production, troubleshoot P1/P2 incidents, perform RCA, validate deployments, test system reliability, and automate operational tasks.
-
Monitor production systems and make sure they are available, fast, and reliable.
-
Build dashboards, alerts, and monitoring to detect issues early.
-
Monitor critical areas such as:
-
APIs
-
DNS
-
Load Balancers
-
WAF
-
MQ
-
Databases / DB2
-
Cloud
-
Telephony
-
Vendor systems
-
-
Create synthetic tests/health checks to make sure important customer and agent journeys are working.
-
Validate systems before and after deployments or major changes.
-
Check DNS, certificates, routing, connectivity, and configurations.
-
Perform load testing, failover testing, recovery testing, and resilience testing.
-
Troubleshoot P1/P2 production incidents and identify the root cause.
-
Conduct RCA/postmortems and make sure the same issue does not happen again.
-
Automate health checks, monitoring, certificate checks, DNS checks, and post-deployment validation.
-
Create and maintain runbooks and troubleshooting tools.
-
Work with application, network, security, cloud, database, telephony, and vendor teams.
-
Strong SRE / Production Engineering / DevOps / Infrastructure experience.
-
Strong production troubleshooting experience.
-
Experience with:
-
Monitoring, logging, metrics, tracing
-
Dashboards and alerting
-
Synthetic monitoring
-
Incident management and RCA
-
Automation
-
-
Good knowledge of:
-
DNS
-
Load Balancers
-
TLS/SSL certificates
-
WAF
-
API Gateways
-
Cloud
-
Service-to-service/API integrations
-
-
Programming/scripting experience in Python, Go, PowerShell, Java, or similar.
-
Experience handling production incidents and performing root cause analysis.
-
Good communication skills and ability to work with multiple technical teams.
-
Salesforce CRM / CTI integrations
-
Contact center / telephony experience
-
Five9
-
Real-time voice or conversational AI
-
IBM MQ
-
DB2 / Mainframe
-
Imperva
-
F5
-
Azure / Google Cloud Platform
-
Performance and capacity testing
-
Chaos/resilience testing
-
Healthcare or other high-availability enterprise environments
Looking forward to qualified submissions only.
- Dice Id: 91135853
- Position Id: 726-38728-1791472830
- Posted 1 hour ago
Company Info
TROR is an artificial intelligence consultancy specializing in developing powerful and customized Al solutions for business. With top Al Experts we take pride in providing the best cutting-edge Al consultancy. Our years of experience in various industries helps us to develop and implement bespoke Al solutions for businesses. Our on demand Al products have helped over 100 companies drive transformational results.
The solutions we bring on your table meet the highest industry standards and quality, effectively and efficiently resolving your issues and optimizing the way you want to move forward in the market. Through our customer centric approach, we ensure that we are always there for our valuable customers by offering them satisfactory solutions for guaranteed results.


Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs