Site Reliability Engineer (SRE)
Full Time
On-site


The Brixton Group
Fitment
Dice Job Match Score™
🔢 Crunching numbers...
Job Details
Skills
- Electronic Commerce
- Akamai
- Performance Engineering
- KPI
- Capacity Management
- Stress Testing
- Performance Analysis
- Backup
- Recovery
- Disaster Recovery
- Computer Networking
- Palo Alto
- Firewall
- Amazon DynamoDB
- API
- SaaS
- Technical Writing
- Confluence
- Production Engineering
- DevOps
- Amazon Web Services
- Virtual Private Cloud
- Amazon CloudFront
- Amazon S3
- WAF
- New Relic
- Incident Management
- Root Cause Analysis
- SLA
- Reliability Engineering
- Apache JMeter
- Load Testing
- Terraform
- Docker
- Git
- Linux
- Unix
- Dragon NaturallySpeaking
- DNS
- TLS
- SSL
- Shell Scripting
- Node.js
- PHP
- Python
- Bash
- Communication
- DICE
Summary
Duration: 12+ Months
Location: 100% Remote
We are seeking a Senior Site Reliability Engineer (SRE) with 10+ years of experience to own the reliability, availability, observability, and production operations of a complex, multi-region global e-commerce platform. The environment is primarily AWS-based, with extensive use of ECS/Docker, Lambda, API Gateway, DynamoDB, CloudFront/Akamai, and SaaS platforms such as VTEX. This is a highly hands-on role that requires deep expertise in production incident management, observability, performance engineering, AWS troubleshooting, and resilience.
Key Responsibilities:
26-01122
#dice
Location: 100% Remote
We are seeking a Senior Site Reliability Engineer (SRE) with 10+ years of experience to own the reliability, availability, observability, and production operations of a complex, multi-region global e-commerce platform. The environment is primarily AWS-based, with extensive use of ECS/Docker, Lambda, API Gateway, DynamoDB, CloudFront/Akamai, and SaaS platforms such as VTEX. This is a highly hands-on role that requires deep expertise in production incident management, observability, performance engineering, AWS troubleshooting, and resilience.
Key Responsibilities:
- Lead P1/P2 production incidents, drive recovery, and conduct RCA and blameless postmortems.
- Own observability using New Relic, AWS CloudWatch, and Amazon Athena.
- Define and monitor SLIs, SLOs, SLAs, reliability KPIs, and alerting strategies.
- Perform capacity planning, load testing, stress testing, and performance analysis using tools such as k6 and JMeter.
- Validate backup/restore, disaster recovery, resilience, patching, runtime upgrades, and SSL/TLS certificates.
- Troubleshoot AWS networking, ALB traffic, CloudFront/CDN, DNS, WAF, DDoS/Bot attacks, and Palo Alto firewall interactions.
- Troubleshoot production issues involving ECS, Lambda, S3, DynamoDB, API Gateway, and VTEX/SaaS integrations.
- Read and assess Terraform/IaC, Docker containers, and production infrastructure for reliability risks.
- Maintain operational runbooks and technical documentation in Confluence.
- 10+ years in SRE, Production Engineering, DevOps, or similar roles.
- Strong hands-on experience with AWS: VPC, ALB, CloudFront, ECS, S3, Lambda, WAF, Secrets Manager.
- Expert-level New Relic, CloudWatch, and Athena experience.
- Strong incident management, RCA, postmortems, SLI/SLO/SLA, and reliability engineering experience.
- Hands-on k6/JMeter performance and load testing.
- Strong knowledge of Terraform, Docker, Git, Linux/Unix, DNS, TLS/SSL, and shell scripting.
- Experience troubleshooting Node.js, PHP, Python, and Bash-based applications.
- Experience with CDN/edge technologies, traffic analysis, security, and production troubleshooting.
- Excellent communication and ability to operate independently in a mission-critical environment.
26-01122
#dice
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 10212185
- Position Id: 26-01122
- Posted 18 hours ago
Company Info
About The Brixton Group
We are a values-based professional technology services firm with over 150 years of combined experience in the technology industry. We re deeply engaged in the successful pairing of the right people to the right projects, and we attribute our national presence to the referrals we ve received through current and past clients and candidates.
Our History
-Established in 1998
-Woman-Owned Business with national footprint
AWARDS
-Four straight years listed on Inc. Magazine s Fastest-Growing Private Companies in America.
-Three straight years listed on Charlotte s Fast 50.
We believe that when people find the ideal setting to express their talents, the possibilities are infinite. This is why we exist. We are inspired and guided by a greater purpose than profit.
Career Opportunities
Our History
-Established in 1998
-Woman-Owned Business with national footprint
AWARDS
-Four straight years listed on Inc. Magazine s Fastest-Growing Private Companies in America.
-Three straight years listed on Charlotte s Fast 50.
We believe that when people find the ideal setting to express their talents, the possibilities are infinite. This is why we exist. We are inspired and guided by a greater purpose than profit.
Career Opportunities
Create job alert
Never miss an opportunity! Create an alert based on the job you applied for.
Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs