SRE - Observability

• Posted 1 day ago • Updated 1 day ago
Full Time
On-site
Fitment

Dice Job Match Score™

📋 Comparing job requirements...

Job Details

Skills

  • Asset Management
  • Scalability
  • Management
  • Capacity Management
  • Performance Tuning
  • Dashboard
  • Instrumentation
  • Build Automation
  • Documentation
  • DevOps
  • Linux Administration
  • Python
  • Scripting
  • Configuration Management
  • Incident Management
  • Problem Solving
  • Conflict Resolution
  • Communication
  • Operational Excellence
  • Software Performance Management
  • Extract
  • Transform
  • Load
  • Financial Services

Summary

Senior Site Reliability Engineer - Observability
Location: Austin, TX Area (Remote-First)
Requirement: Candidates must be within commuting distance of Austin.

About the Role
A leading asset management firm is seeking a Senior Site Reliability Engineer to own and enhance its observability platforms. This role combines operational support with platform engineering, focused on improving reliability, scalability, and visibility across the technology organization.

Responsibilities
* Own the availability, performance, and support of observability platforms.
* Serve as an escalation point for monitoring and logging-related issues.
* Manage platform upgrades, capacity planning, performance tuning, and operational health.
* Partner with engineering teams to improve dashboards, alerting, and instrumentation.
* Build automation and self-service capabilities that reduce operational toil.
* Develop and maintain infrastructure-as-code and platform standards.
* Drive observability best practices and platform modernization initiatives.
* Maintain documentation, runbooks, and incident response processes.

Requirements
* 5+ years of experience in SRE, DevOps, Platform Engineering, or related roles.
* Strong hands-on experience with enterprise observability, logging, and monitoring platforms.
* Experience operating large-scale on-premises infrastructure and distributed systems.
* Strong Linux administration and troubleshooting skills.
* Proficiency in Python and scripting for automation.
* Experience with infrastructure automation and configuration management tools.
* Proven incident management and problem-solving capabilities.
* Strong communication skills and a passion for operational excellence.

Preferred Experience
* Modern observability and APM platforms.
* Distributed tracing and telemetry frameworks.
* Log collection and data pipeline technologies.
* Infrastructure-as-code and automation tooling.
* Financial services or other highly regulated environments.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90922487
  • Position Id: 24595566
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Austin, Texas

Today

Full-time

USD 130,000.00 - 175,000.00 per year

Austin, Texas

Today

Full-time

USD 135,200.00 - 306,400.00 per year

Austin, Texas

Today

Full-time

USD 88,000.00 - 136,900.00 per year

New Jersey

Today

Full-time

USD 149,000.00 - 186,000.00 per year

Search all similar jobs