Site Reliability Engineer
Veterans Affairs ESOM
Overview
We are seeking an experienced Site Reliability Engineer (SRE) to support enterprise Problem Management and Root Cause Analysis (RCA) activities. In this senior-level role, you will investigate complex production outages and incidents across a large, diverse environment, identifying technical root causes and supporting corrective actions. This opportunity is ideal for a technically versatile professional comfortable working across multiple infrastructure and application domains to ensure system reliability and sustainable resolution.
Education Requirements
- Bachelor's Degree in Engineering, Computer Science, Systems, Business, or a related scientific/technical discipline.
Certification Requirements
There are no certification requirements for this role.
Clearance Requirements
Clearance Level: Public Trust
Work Arrangement
Remote
Responsibilities
- Investigate complex production outages and major incidents using logs, telemetry, monitoring data, traces, incident history, change records, and other technical evidence.
- Assess unfamiliar systems to determine the information needed for effective technical investigations.
- Query, extract, correlate, and analyze data using enterprise observability platforms and scripting/query languages.
- Develop and test root cause hypotheses, challenge unsupported conclusions, and validate findings against technical evidence.
- Develop technical mitigation and corrective action recommendations focused on sustainable resolution.
- Support verification that proposed fixes address the original failure conditions through testing, simulation, or other defensible methods.
- Identify recurring patterns and systemic risks across systems and teams.
- Collaborate with technical teams, investigation coordinators, analysts, and stakeholders throughout RCA activities.
- Leverage approved AI-assisted tools to accelerate analysis while independently validating results.
Required Qualifications
- Demonstrated experience diagnosing complex, enterprise-scale production outages across multiple technology domains.
- Hands-on experience querying enterprise observability and log analysis platforms such as Splunk, Dynatrace, Elastic/ELK, Grafana, or Datadog.
- Proficiency with scripting and query languages such as SQL, Python, or PowerShell.
- Ability to correlate logs, metrics, changes, incidents, and operational data into clear technical timelines and evidence-based conclusions.
- Strong analytical judgment with the ability to independently evaluate, validate, or challenge proposed root causes.
- Excellent communication and collaboration skills across engineering, support, and technical teams.
Desired Skills
- Experience working with enterprise observability and monitoring platforms.
- Ability to work across diverse technical environments and analyze telemetry data effectively.
- Strong problem-solving skills and attention to detail.
- Proven ability to work independently and handle complex investigations.
Pay Range
Pay range: $80,000.00–$110,000.00 (annual). This pay range is based on qualifications and experience, and assumes all role requirements are met.
Why Apply
This is a unique opportunity to contribute to enterprise-wide system stability and reliability through complex technical investigations. If you thrive in challenging environments and enjoy solving complex production issues, we encourage you to apply and join a dedicated team committed to operational excellence.