Site Reliability Engineer - (GPS-Salesforce CRM & YAVA)

Hybrid in Hartford, CT, US β€’ Posted 1 hour ago β€’ Updated 1 hour ago
Contract W2
Contract Independent
12 Months
Hybrid
Depends on Experience
Fitment

Dice Job Match Scoreβ„’

πŸ”’ Crunching numbers...

Job Details

Skills

  • Site Reliability Engineer
  • Reduce P1/P2 incidents
  • Proactive Monitoring and Early Detection
  • infrastructure dependencies.
  • Hands-on experience with metrics
  • logs
  • traces
  • dashboards
  • alerting
  • and synthetic monitoring.
  • Python
  • Go
  • PowerShell
  • Java
  • or similar languages.
  • root cause analysis
  • postmortems
  • change validation
  • and automation.
  • Salesforce CRM and CTI integrations.

Summary

Job Title: Site Reliability Engineer - (GPS-Salesforce CRM & YAVA)

Location: Hartford, CT (Hybrid)

Type : (Contract)

Hands-on Reliability Engineering Role

PRIMARY OUTCOME: Prevent incidents

PLATFORM SCOPE: GPS and YAVA

ROLE TYPE: Hands-on engineering

Role Summary

The Site Reliability Engineer is responsible for preventing production incidents across the GPS Salesforce CRM and YAVA servicing platforms. This hands-on role identifies reliability risks, builds proactive detection and automation, validates high-risk changes, tests resilience, and converts incident lessons into permanent engineering improvements.

Primary Outcome

Reduce P1/P2 incidents, detect degradation before customers or agents are affected, restore service faster when failures occur, and prevent the same failure from recurring.

Core Responsibilities

  1. Reliability Indicators and Service Health
  • Implement and maintain SLIs, SLOs, and dashboards for incident-relevant platform behavior.
  • For GPS, monitor availability, login/authentication, member-search success, CTI or screen-pop success, critical API success, and p95/p99 response time.
  • For YAVA, monitor call/session setup, critical transaction success, escalation-to-agent success, API dependency health, and end-to-end latency.
  • Measure shared dependency health across DNS, load balancing, WAF, API gateway, MQ, mainframe/DB2, cloud, telephony, and vendors.
  1. Proactive Monitoring and Early Detection
  • Build synthetic probes for critical GPS agent and YAVA customer journeys using safe production validation patterns.
  • Create dashboards and alerts that identify the failing hop and correlate latency, errors, traffic, and concurrency.
  • Implement leading indicators for configuration drift, rising queue depth, DNS failures, certificate risk, routing mismatches, capacity saturation, and tail latency.
  • Validate what is actually served in production, including live certificates, DNS answers, routing, and load balancer pool membership.
  • Tune alerts so each alert is actionable, correctly routed, and linked to a runbook.
  1. Change and Upgrade Validation
  • Assess technical risk for in-scope application, infrastructure, platform, security, and vendor changes.
  • Define and execute pre-change checks, post-change verification, and next-business-peak validation when impact may be delayed.
  • Verify that high-risk changes include regression evidence, rollback criteria, and tested recovery procedures.
  • Correlate alerts and incidents with the change timeline to distinguish planned work from unexpected regression.
  • Automate repeatable validation and provide self-service checks to teams making changes.
  1. Dependency and Configuration Reliability
  • Maintain current technical dependency maps, owners, escalation paths, and critical configuration assumptions.
  • Continuously validate DNS, certificates, firewall/WAF behavior, load balancer health checks, API routing, MQ connectivity, database connectivity, and vendor endpoints.
  • Identify missing monitoring, latent configuration defects, unsupported components, and single points of failure.
  • Track platform lifecycle and certificate risks and drive remediation before they create incidents.
  1. Resilience, Capacity, and Failure Containment
  • Run load, failover, recovery, and targeted chaos tests for critical paths.
  • Assess capacity against call volume, agent demand, transaction volume, concurrency, MQ limits, connection pools, and backend constraints.
  • Review retries, timeouts, circuit breakers, and graceful degradation to prevent cascading failures and backlogs.
  • Validate fallback options, including transfer or alternate routing when a dependent platform or vendor is unavailable.
  1. Incident Analysis and Prevention
  • Support major-incident diagnosis with end-to-end telemetry and dependency knowledge.
  • Lead or contribute to blameless technical postmortems and evidence-based root cause analysis.
  • Ensure preventive actions are specific, testable, owned, and time-bound.
  • Analyze trends by change source, dependency, detection method, impact duration, and repeat failure mode.
  • Implement or coordinate permanent fixes and verify that corrective actions reduce the identified risk.
  1. Automation and Toil Reduction
  • Develop automation for health checks, configuration validation, certificate checks, DNS and routing tests, post-change verification, and recurring diagnostics.
  • Create reusable tooling and runbooks that reduce manual error and speed triage.
  • Eliminate recurring operational work when engineering automation provides a safer and repeatable solution.
  1. Cross-Team Execution
  • Partner with GPS, YAVA, application, network, security, telephony, MQ/mainframe, database, cloud, and vendor teams.
  • Communicate risks, findings, and recommendations clearly to both technical and leadership audiences.
  • Drive actions through influence and evidence while respecting ownership of the underlying platforms.
  • Participate in production-readiness, change, incident, and reliability review forums.

Success Measures

Measure Expected Direction

Incident reduction Fewer P1/P2 and customer- or agent-impacting incidents

Repeat failures Fewer recurrences of known failure modes

Detection More degradation detected before external reports

MTTD and MTTR Faster detection, diagnosis, and restoration

Validation coverage More high-risk changes verified before and after implementation

Observability coverage More critical GPS and YAVA journeys monitored end to end

Automation Less manual validation and operational toil

Corrective actions Higher completion and verified effectiveness

Required Qualifications

  • Experience in Site Reliability Engineering, production engineering, DevOps, platform engineering, infrastructure engineering, or large-scale production operations.
  • Strong troubleshooting skills across distributed applications and shared infrastructure dependencies.
  • Hands-on experience with metrics, logs, traces, dashboards, alerting, and synthetic monitoring.
  • Knowledge of networking, DNS, load balancing, TLS/certificates, WAFs, API gateways, cloud services, and service-to-service integrations.
  • Programming or scripting experience with Python, Go, PowerShell, Java, or similar languages.
  • Experience with incident response, root cause analysis, postmortems, change validation, and automation.
  • Ability to communicate clearly and influence teams outside the direct reporting line.

Preferred Experience

  • Salesforce CRM and CTI integrations.
  • Contact-center platforms, telephony, Five9, real-time voice, or conversational AI.
  • IBM MQ, DB2, mainframe integrations, Imperva, F5, Azure, or Google Cloud.
  • Performance engineering, capacity testing, chaos engineering, and resilience testing.
  • Regulated healthcare or another high-availability enterprise environment.
Employers have access to artificial intelligence language tools (β€œAI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10200946b
  • Position Id: 9109368
  • Posted 1 hour ago
Contact the job poster
RR

Rahul Rawat

Recruiter @ Nityo Infotech Corporation
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hartford, Connecticut

β€’

2d ago

Easy Apply

Contract, Third Party

Depends on Experience

Hartford, Connecticut

β€’

Today

Easy Apply

Full-time

$55 - $65 per hour

Hartford, Connecticut

β€’

2d ago

Easy Apply

Third Party, Contract

$60 - $70

Remote

β€’

Today

Full-time

USD 100,000.00 - 120,000.00 per year

Search all similar jobs