Senior Site Reliability Engineer

Southlake, TX, US • Posted 6 hours ago • Updated 6 hours ago
Full Time
On-site
USD $130,000.00 - 155,000.00 per year
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Creative Problem Solving
  • Finance
  • Operational Excellence
  • Cryptography
  • Prism
  • Continuous Improvement
  • Decision-making
  • Operational Efficiency
  • Application Development
  • Adaptability
  • Software Engineering
  • Reliability Engineering
  • DevOps
  • Management
  • Service Level
  • Health Care Administration
  • Splunk
  • Grafana
  • AppDynamics
  • ACH
  • .NET
  • Java
  • Artificial Intelligence
  • Productivity
  • GitHub
  • Microsoft
  • Continuous Integration
  • Continuous Delivery
  • JIRA
  • Bamboo
  • Confluence
  • BMC Control-M
  • Microsoft SQL Server
  • IBM DB2
  • IT Service Management
  • Analytical Skill
  • Communication
  • Collaboration
  • Conflict Resolution
  • Problem Solving
  • Pivotal
  • Cloud Foundry
  • Google Cloud Platform
  • Google Cloud
  • Linux Administration
  • Red Hat Enterprise Linux
  • Shell Scripting
  • Python
  • Perl
  • Ruby
  • Windows PowerShell
  • Disaster Recovery
  • Business Continuity Planning
  • Financial Services
  • Network
  • Dragon NaturallySpeaking
  • DNS
  • Virtual Private Network
  • Proxies
  • Firewall
  • Storage
  • Mentorship

Summary

Your Opportunity

At Schwab, you're empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together. We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).

As a Site Reliability Engineer supporting the Cashiering organization, you will play a critical role in ensuring the stability, resiliency, and operational excellence of Schwab's Move Money platforms. This role supports business-critical modernization initiatives, including Crypto and PRISM, while driving production readiness, availability, and continuous improvement across distributed applications. You will leverage observability, automation, and data-driven decision making to identify issues, improve system reliability, and optimize operational efficiency.

Working closely with application development teams, infrastructure partners, business stakeholders, and external vendors, you will lead complex incident resolution efforts, influence resiliency strategies, and help shape operational best practices. Success in this role requires strong problem-solving skills, adaptability in a fast-paced environment, and the ability to build collaborative relationships while balancing reliability, performance, and business outcomes across mission-critical Cashiering services.

What you have

Required Qualifications
  • 5+ years of experience supporting production applications, software engineering, site reliability engineering, DevOps, or related technology environments
  • Experience managing application reliability through Service Level Objectives (SLOs), monitoring practices, and operational health management
  • Demonstrated ability to troubleshoot and resolve complex issues across applications, platforms, infrastructure, and distributed systems
  • Experience with observability and monitoring tools including Splunk, Grafana, Datadog, AppDynamics, or ThousandEyes
  • Strong knowledge of Cashiering and Move Money functions including Wires, ACH, Journals, Checks, and Check Deposits
  • Experience reviewing and troubleshooting .NET and Java-based applications
  • Experience with AI-enabled productivity and automation tools such as GitHub Copilot and Microsoft Copilot Studio
  • Experience supporting CI/CD pipelines, production readiness processes, and operational support practices
  • Working knowledge of Jira, Bamboo, Confluence, Control-M, SQL Server, DB2, and IT service management disciplines
  • Strong analytical thinking, communication, collaboration, and problem-solving skills with the ability to coordinate effectively during critical incidents

Preferred Qualifications
  • Experience with Pivotal Cloud Foundry and Google Cloud Platform environments
  • Knowledge of Linux administration, including RHEL environments
  • Experience with shell scripting, Python, Perl, Ruby, or PowerShell
  • Understanding of disaster recovery, resiliency planning, and business continuity programs
  • Experience supporting distributed applications in highly regulated or financial services environments
  • Knowledge of network services including DNS, VPNs, load balancers, proxies, and firewalls
  • Experience working with SAN and NAS storage technologies
  • Experience mentoring team members and promoting operational standards and best practices

In addition to the salary range, this role is eligible for bonuses or incentive opportunity
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90989465
  • Position Id: 67c164e745b42b5fb706a5c12c81172
  • Posted 6 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Coppell, Texas

Today

Easy Apply

Full-time

$40 - $48 per hour

Southlake, Texas

Today

Contract

USD 50.00 - 55.00 per hour

Plano, Texas

Today

Full-time

Plano, Texas

Today

Full-time

Search all similar jobs