The ideal candidate will be a hands-on infrastructure/platform engineer with strong experience in AWS, Terraform, Linux, networking, containers, automation, and production reliability. The role will involve infrastructure delivery, automation, troubleshooting, incident response, and production support, along with supporting data platform workloads.
Design, implement, maintain, and troubleshoot AWS cloud infrastructure across production environments.
Manage AWS services including VPC, IAM, ECS/Fargate, Lambda, S3, and RDS.
Develop and maintain reusable Terraform modules, remote state, environment separation, and infrastructure-as-code practices.
Support infrastructure changes through CI/CD pipelines, code reviews, and safe deployment/rollback procedures.
Monitor production environments, respond to incidents, perform root cause analysis, and implement permanent fixes.
Troubleshoot Linux systems, DNS, TLS, routing, load balancing, connectivity, and resource utilization issues.
Manage containerized workloads using Docker and AWS ECS/Fargate.
Develop automation using Python and Shell scripting to reduce manual operational activities.
Implement and maintain security, access management, secrets management, backup, recovery, and disaster recovery practices.
Perform capacity planning, resource optimization, and AWS cost optimization.
Support production data platforms and troubleshoot issues involving Databricks/Redshift, Airflow, dbt, and SQL-based workloads.
Investigate data pipeline failures, data freshness issues, dependencies, retries, backfills, and recovery procedures.
Collaborate with engineering and data teams to improve platform reliability, scalability, and operational efficiency.
9+ years of experience in Cloud Infrastructure, SRE, Platform Engineering, DevOps, or a related field.
Strong hands-on experience with AWS in production environments.
Strong Terraform experience, including modules, remote state, and environment management.
Strong knowledge of Linux and cloud networking.
Experience with Docker and containerized environments.
Experience with Python and/or Shell scripting for automation.
Strong understanding of IAM, security, secrets management, backups, and disaster recovery.
Experience with production support, incident response, troubleshooting, and root cause analysis.
Experience with monitoring, alerting, reliability, capacity planning, and performance optimization.
Experience supporting production data workloads with SQL and data operations.