Sr SRE Automation Engineer

Austin, TX, US • Posted 6 hours ago • Updated 18 minutes ago
Contract W2
On-site
DOE
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Operational Excellence
  • Software Design
  • Management
  • Scrum
  • Performance Monitoring
  • Dashboard
  • Instrumentation
  • Analytics
  • Workflow
  • Scalability
  • Operational Efficiency
  • Kubernetes
  • Amazon Web Services
  • Splunk
  • DevOps
  • Scripting
  • Cloud Computing
  • High Availability
  • Artificial Intelligence
  • Machine Learning (ML)
  • Continuous Integration
  • Continuous Delivery
  • Computer Networking
  • System Administration
  • Orchestration
  • Identity Management
  • Reliability Engineering
  • Incident Management
  • Production Support
  • Root Cause Analysis
  • Capacity Management
  • Forecasting
  • Documentation
  • Communication
  • Collaboration
  • Agile
  • Financial Services

Summary

Key Responsibilities
Champion Site Reliability Engineering (SRE) principles and promote automation-driven operational excellence
Identify opportunities to build innovative tools and solutions that address complex operational challenges across enterprise and mission-critical applications
Design, develop, and maintain automation solutions that reduce manual effort and improve operational efficiency
Create scripts and automation frameworks to streamline infrastructure management, deployment processes, and operational workflows
Design and implement AI/ML-driven automation pipelines, anomaly detection, predictive alerting, and intelligent operational response solutions
Enhance observability capabilities through advanced monitoring, telemetry, logging, and analytics platforms
Lead the expansion of automation coverage across deployment, monitoring, alerting, remediation, and self-healing workflows
Collaborate with Engineering, Scrum, Operations, and Infrastructure teams to improve system availability, reliability, and performance
Monitor, triage, troubleshoot, and resolve critical production incidents and platform issues
Implement operational changes with minimal risk while ensuring effective stakeholder communication
Develop tools, frameworks, dashboards, and instrumentation to improve application deployment success and operational visibility
Leverage AI/ML capabilities to enhance platform monitoring, rollout validation, and operational intelligence
Drive adoption of AIOps platforms and machine learning-assisted observability practices
Support capacity planning and performance forecasting using data-driven analytics and predictive models
Design and implement CI/CD orchestration solutions to accelerate software delivery and improve deployment reliability
Promote GitOps methodologies and automation best practices across engineering teams
Troubleshoot mission-critical application workflows and collaborate with development teams to address reliability concerns
Develop and maintain operational runbooks, support procedures, and knowledge documentation
Participate in on-call support rotations and incident response activities
Continuously identify opportunities to improve platform scalability, resilience, performance, and operational efficiency

Required Qualifications
7+ years of experience
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience
Certified Kubernetes Administrator (CKA)
AWS Certified DevOps Engineer
AWS Certified SysOps Administrator
Google Professional Cloud DevOps Engineer
Splunk Certification
Relevant Cloud, DevOps, SRE, AIOps, or Observability certifications

Skills
Automation Frameworks and Operational Tooling
Scripting and Programming
Cloud Platforms
Distributed Systems and High-Availability Architectures
Monitoring, Logging, Observability, and Alerting Solutions
AIOps and ML-Assisted Observability
AI/ML-Driven Operational Automation
CI/CD Pipelines and Deployment Automation
GitOps
Networking, Infrastructure, and System Administration
Container Orchestration
Monitoring Tools
Identity and Access Management Platforms
Site Reliability Engineering (SRE) Practices
Incident Management
Production Support
Root Cause Analysis
Capacity Planning and Performance Forecasting
Operational Runbooks and Knowledge Documentation
Stakeholder Communication
Cross-functional Collaboration
Agile Delivery
Financial Services Industry Experience

Schedule
Start date: 2026-09-22
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: compun
  • Position Id: SAHDC5895784
  • Posted 6 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Austin, Texas

Today

Contract

USD 75.00 - 80.00 per hour

Hybrid in Austin, Texas

2d ago

Easy Apply

Contract

Depends on Experience

Austin, Texas

Today

Easy Apply

Contract

$80 - $90

Austin, Texas

Today

Full-time

Search all similar jobs