Join a growing Site Reliability Engineering team focused on building dependable, scalable digital platforms. This role combines cloud infrastructure, automation, monitoring, and software development to support digital transformation initiatives and high-traffic online applications. You'll partner directly with product and operations teams to improve performance, reliability, and delivery across the organization's technology environment.
This is an opportunity for an engineer who enjoys both building systems and helping teams work better. You'll have a direct hand in creating automation, shaping infrastructure standards, improving observability, and making sure applications are ready to scale. If you're comfortable exploring new technologies, solving complex reliability challenges, and taking ownership of projects from design through deployment, you'll have plenty of room to make an impact.
Required Skills & Experience
3+ years of experience coding in Python or Ruby
2+ years of experience with AWS technologies
3+ years of experience supporting business processes or technology implementations
Hands-on experience with Chef, Puppet, Ansible, or similar configuration-management tools
Experience scaling cloud-based applications with a focus on automation and reliability
Bachelor's or advanced degree in Information Technology, Computer Science, Business, or a related field-or equivalent experience
Desired Skills & Experience
2+ years of experience leading web technology or application delivery initiatives
Experience supporting eCommerce or high-volume online applications
Familiarity with continuous integration, continuous delivery, and release automation
Experience with Dynatrace, Splunk, Rigor, Quantum, or similar monitoring platforms
Experience centralizing application logs with Splunk
Strong understanding of application performance, infrastructure design, and operational support
What You Will Be Doing
Lead infrastructure, automation, and reliability efforts for digital transformation initiatives
Define infrastructure and system requirements using SRE and cloud best practices
Build automation that helps applications and environments scale efficiently
Design and implement monitoring, alerting, and observability solutions
Partner with product and operations teams to address reliability and performance needs
Manage application infrastructure, release processes, performance metrics, and platform improvements
Tech Breakdown
30% Cloud Infrastructure and Application Scalability
25% Automation and Configuration Management
20% Monitoring, Observability, and Log Management
15% Software Development and Scripting
10% Release Management and Cross-Functional Collaboration
Daily Responsibilities
Build and maintain scalable cloud infrastructure for online applications
Develop automation using Python, Ruby, Chef, Puppet, Ansible, or similar tools
Create and improve monitoring, alerting, and logging solutions
Review application performance and identify opportunities to improve reliability
Partner with product and operations teams on infrastructure and SRE needs
Support release planning, deployment activities, and continuous delivery improvements
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 10105282
- Position Id: 887350
- Posted 13 hours ago