Site Reliability Engineer (SRE) Mobile & Digital Observability

• Posted 2 days ago • Updated 54 minutes ago
Contract Independent
Contract Corp To Corp
Contract W2
$57/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • SRE
  • Mobile Observability
  • Android
  • IOS
  • SignalFx
  • Splunk

Summary

Position: Site Reliability Engineer (SRE) Mobile & Digital Observability
Location: Atlanta, GA (Onsite)

Role Overview:
The Senior Site Reliability Engineer (SRE) is a hands-on role responsible for the availability, performance, and end-to-end observability of QSR digital platforms across Mobile (iOS/Android), Web, and POS systems.

This role is part of the Observability team and works closely with mobile, web, and backend engineering teams to ensure full visibility into customer journeys and user experience.

The focus is on building and operating Real User Monitoring (RUM), synthetic monitoring, and end-to-end telemetry correlation using Splunk and SignalFx, ensuring issues are detected before customer impact. This is not a monitoring-only role-it requires active involvement in instrumentation, release observability, and reliability engineering.
Key Responsibilities
- Define and enforce SLIs, SLOs, and error budgets for critical customer journeys (ordering, checkout, payments)
- Own end-to-end observability across Mobile, Web, and POS platforms
- Implement and operate RUM and synthetic monitoring for customer-facing journeys
- Build mobile-first monitoring coverage including app performance, crash rates, API performance, and user journey tracking
- Use Splunk and SignalFx to design dashboards, detectors, and actionable alerts
- Enable correlation across mobile CDN API backend systems using logs, metrics, and traces
- Partner with engineering teams for instrumentation, SDK integration, and embedding observability into releases
- Analyze telemetry to detect post-release issues, device/OS-specific failures, and network degradation
- Lead response for P1/P2 incidents and drive root cause analysis
- Automate operational toil and improve reliability
Required Skills
- Strong experience with Splunk (logs, dashboards) and SignalFx (metrics/APM)
- Hands-on with RUM and synthetic monitoring tools
- Experience with mobile observability (iOS/Android), including performance monitoring and crash analysis
- Strong understanding of distributed systems and microservices (Java, Node.js)
- Experience with Azure (AKS, App Services, APIM)
- Ability to correlate frontend issues with backend services
- Experience with CI/CD pipelines and observability in release processes
What Success Looks Like
- Clear visibility into mobile, web, and POS customer journeys
- Issues identified before customer complaints or app store feedback
- Strong, low-noise user-impact-driven alerting
- Reduced crash rates, latency, and checkout failures
- Observability embedded into every release

Thanks & Regards
Amarjit


Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91163888
  • Position Id: 2026-78
  • Posted 2 days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

•

15d ago

Easy Apply

Contract, Third Party

$63 - $65

Remote

•

13d ago

Easy Apply

Contract, Third Party

70 - 75

No location provided

•

Today

Easy Apply

Full-time, Contract, Third Party

$DOE

Dallas, Texas

•

Today

Easy Apply

Contract

Depends on Experience

Search all similar jobs