Key Responsibilities
Champion Site Reliability Engineering (SRE) principles and promote automation-driven operational excellence
Identify opportunities to build innovative tools and solutions that address complex operational challenges across enterprise and mission-critical applications
Design, develop, and maintain automation solutions that reduce manual effort and improve operational efficiency
Create scripts and automation frameworks to streamline infrastructure management, deployment processes, and operational workflows
Design and implement AI/ML-driven automation pipelines, anomaly detection, predictive alerting, and intelligent operational response solutions
Enhance observability capabilities through advanced monitoring, telemetry, logging, and analytics platforms
Lead the expansion of automation coverage across deployment, monitoring, alerting, remediation, and self-healing workflows
Collaborate with Engineering, Scrum, Operations, and Infrastructure teams to improve system availability, reliability, and performance
Monitor, triage, troubleshoot, and resolve critical production incidents and platform issues
Implement operational changes with minimal risk while ensuring effective stakeholder communication
Develop tools, frameworks, dashboards, and instrumentation to improve application deployment success and operational visibility
Leverage AI/ML capabilities to enhance platform monitoring, rollout validation, and operational intelligence
Drive adoption of AIOps platforms and machine learning-assisted observability practices
Support capacity planning and performance forecasting using data-driven analytics and predictive models
Design and implement CI/CD orchestration solutions to accelerate software delivery and improve deployment reliability
Promote GitOps methodologies and automation best practices across engineering teams
Troubleshoot mission-critical application workflows and collaborate with development teams to address reliability concerns
Develop and maintain operational runbooks, support procedures, and knowledge documentation
Participate in on-call support rotations and incident response activities
Continuously identify opportunities to improve platform scalability, resilience, performance, and operational efficiency
Required Qualifications
7+ years of experience
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience
Certified Kubernetes Administrator (CKA)
AWS Certified DevOps Engineer
AWS Certified SysOps Administrator
Google Professional Cloud DevOps Engineer
Splunk Certification
Relevant Cloud, DevOps, SRE, AIOps, or Observability certifications
Skills
Automation Frameworks and Operational Tooling
Scripting and Programming
Cloud Platforms
Distributed Systems and High-Availability Architectures
Monitoring, Logging, Observability, and Alerting Solutions
AIOps and ML-Assisted Observability
AI/ML-Driven Operational Automation
CI/CD Pipelines and Deployment Automation
GitOps
Networking, Infrastructure, and System Administration
Container Orchestration
Monitoring Tools
Identity and Access Management Platforms
Site Reliability Engineering (SRE) Practices
Incident Management
Production Support
Root Cause Analysis
Capacity Planning and Performance Forecasting
Operational Runbooks and Knowledge Documentation
Stakeholder Communication
Cross-functional Collaboration
Agile Delivery
Financial Services Industry Experience
Schedule
Start date: 2026-09-22
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: compun
- Position Id: SAHDC5895784
- Posted 6 hours ago