Role: Data Infrastructure Site Reliability Engineer (SRE) – AWS & Big Data Platforms (Apple)
Location: Dallas, Texas (Remote)
Contract to hire after 3 months
Duration: 12 Months
Skills - Site Reliability Engineering, Big Data, AWS, AWS EMR, AWS EKS, AWS MSK, AWS Athena, AWS Glue, Spark, Iceberg, Python, Terraform, Cloudera Hadoop, Agentic Ai, Monitoring Tools
Job Description:
Role Overview
We are seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-scale data infrastructure environments across AWS and on-premises Hadoop platforms. This role requires a strong reliability engineering mindset focused on platform stability, performance, observability, automation, and incident response.
As a key member of the Data Infrastructure SRE team, you will ensure the availability, scalability, security, and operational excellence of mission-critical data platforms while driving continuous improvements through Infrastructure as Code (IaC), AI-enabled automation, and modern SRE practices.
Key Responsibilities
- Maintain and support highly available, scalable, and secure data infrastructure platforms across AWS and on-premises environments.
- Drive operational excellence through automation of repetitive tasks, incident reduction, and proactive reliability improvements.
- Monitor platform health, troubleshoot complex issues, and lead root cause analysis efforts to minimize downtime and improve system resiliency.
- Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives.
- Participate in on-call rotations and partner with teams across the US and India to provide 24x7 operational support. US support is aligned to Pacific Time zone.
- Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise.
Required Skills & Experience
Site Reliability Engineering (SRE)
- Strong SRE mindset with a proven focus on reliability, availability, performance optimization, incident management, and operational excellence.
- Experience delivering services within defined SLAs and ensuring timely resolution of production issues.
- Expertise in troubleshooting complex distributed systems and identifying root causes quickly and effectively.
- AWS & Cloud Infrastructure
- Deep hands-on experience with AWS services, including:
- EMR
- EKS
- MSK
- Athena
- Glue
- IAM
- Amazon S3
- VPC
- AWS networking and security services
- Strong understanding of cloud-native architectures, scalability, and infrastructure resilience.