Job Title: Site Reliability Engineer - (GPS-Salesforce CRM & YAVA)
Location: Hartford, CT (Hybrid)
Type : (Contract)
Hands-on Reliability Engineering Role
PRIMARY OUTCOME: Prevent incidents
PLATFORM SCOPE: GPS and YAVA
ROLE TYPE: Hands-on engineering
Role Summary
The Site Reliability Engineer is responsible for preventing production incidents across the GPS Salesforce CRM and YAVA servicing platforms. This hands-on role identifies reliability risks, builds proactive detection and automation, validates high-risk changes, tests resilience, and converts incident lessons into permanent engineering improvements.
Primary Outcome
Reduce P1/P2 incidents, detect degradation before customers or agents are affected, restore service faster when failures occur, and prevent the same failure from recurring.
Core Responsibilities
- Reliability Indicators and Service Health
- Implement and maintain SLIs, SLOs, and dashboards for incident-relevant platform behavior.
- For GPS, monitor availability, login/authentication, member-search success, CTI or screen-pop success, critical API success, and p95/p99 response time.
- For YAVA, monitor call/session setup, critical transaction success, escalation-to-agent success, API dependency health, and end-to-end latency.
- Measure shared dependency health across DNS, load balancing, WAF, API gateway, MQ, mainframe/DB2, cloud, telephony, and vendors.
- Proactive Monitoring and Early Detection
- Build synthetic probes for critical GPS agent and YAVA customer journeys using safe production validation patterns.
- Create dashboards and alerts that identify the failing hop and correlate latency, errors, traffic, and concurrency.
- Implement leading indicators for configuration drift, rising queue depth, DNS failures, certificate risk, routing mismatches, capacity saturation, and tail latency.
- Validate what is actually served in production, including live certificates, DNS answers, routing, and load balancer pool membership.
- Tune alerts so each alert is actionable, correctly routed, and linked to a runbook.
- Change and Upgrade Validation
- Assess technical risk for in-scope application, infrastructure, platform, security, and vendor changes.
- Define and execute pre-change checks, post-change verification, and next-business-peak validation when impact may be delayed.
- Verify that high-risk changes include regression evidence, rollback criteria, and tested recovery procedures.
- Correlate alerts and incidents with the change timeline to distinguish planned work from unexpected regression.
- Automate repeatable validation and provide self-service checks to teams making changes.
- Dependency and Configuration Reliability
- Maintain current technical dependency maps, owners, escalation paths, and critical configuration assumptions.
- Continuously validate DNS, certificates, firewall/WAF behavior, load balancer health checks, API routing, MQ connectivity, database connectivity, and vendor endpoints.
- Identify missing monitoring, latent configuration defects, unsupported components, and single points of failure.
- Track platform lifecycle and certificate risks and drive remediation before they create incidents.
- Resilience, Capacity, and Failure Containment
- Run load, failover, recovery, and targeted chaos tests for critical paths.
- Assess capacity against call volume, agent demand, transaction volume, concurrency, MQ limits, connection pools, and backend constraints.
- Review retries, timeouts, circuit breakers, and graceful degradation to prevent cascading failures and backlogs.
- Validate fallback options, including transfer or alternate routing when a dependent platform or vendor is unavailable.
- Incident Analysis and Prevention
- Support major-incident diagnosis with end-to-end telemetry and dependency knowledge.
- Lead or contribute to blameless technical postmortems and evidence-based root cause analysis.
- Ensure preventive actions are specific, testable, owned, and time-bound.
- Analyze trends by change source, dependency, detection method, impact duration, and repeat failure mode.
- Implement or coordinate permanent fixes and verify that corrective actions reduce the identified risk.
- Automation and Toil Reduction
- Develop automation for health checks, configuration validation, certificate checks, DNS and routing tests, post-change verification, and recurring diagnostics.
- Create reusable tooling and runbooks that reduce manual error and speed triage.
- Eliminate recurring operational work when engineering automation provides a safer and repeatable solution.
- Cross-Team Execution
- Partner with GPS, YAVA, application, network, security, telephony, MQ/mainframe, database, cloud, and vendor teams.
- Communicate risks, findings, and recommendations clearly to both technical and leadership audiences.
- Drive actions through influence and evidence while respecting ownership of the underlying platforms.
- Participate in production-readiness, change, incident, and reliability review forums.
Success Measures
Measure Expected Direction
Incident reduction Fewer P1/P2 and customer- or agent-impacting incidents
Repeat failures Fewer recurrences of known failure modes
Detection More degradation detected before external reports
MTTD and MTTR Faster detection, diagnosis, and restoration
Validation coverage More high-risk changes verified before and after implementation
Observability coverage More critical GPS and YAVA journeys monitored end to end
Automation Less manual validation and operational toil
Corrective actions Higher completion and verified effectiveness
Required Qualifications
- Experience in Site Reliability Engineering, production engineering, DevOps, platform engineering, infrastructure engineering, or large-scale production operations.
- Strong troubleshooting skills across distributed applications and shared infrastructure dependencies.
- Hands-on experience with metrics, logs, traces, dashboards, alerting, and synthetic monitoring.
- Knowledge of networking, DNS, load balancing, TLS/certificates, WAFs, API gateways, cloud services, and service-to-service integrations.
- Programming or scripting experience with Python, Go, PowerShell, Java, or similar languages.
- Experience with incident response, root cause analysis, postmortems, change validation, and automation.
- Ability to communicate clearly and influence teams outside the direct reporting line.
Preferred Experience
- Salesforce CRM and CTI integrations.
- Contact-center platforms, telephony, Five9, real-time voice, or conversational AI.
- IBM MQ, DB2, mainframe integrations, Imperva, F5, Azure, or Google Cloud.
- Performance engineering, capacity testing, chaos engineering, and resilience testing.
- Regulated healthcare or another high-availability enterprise environment.