Amazon s RFM team is seeking an engineer to optimize DRAM usage across EMR on EC2 clusters through cluster tuning, Spark/YARN optimization, and EC2 instance right-sizing. The role involves analyzing utilization, identifying over-provisioned resources, running optimization experiments, and delivering measurable memory/cost savings without impacting performance or SLAs.
Key Responsibilities
Analyze EMR, Spark, YARN, and CloudWatch metrics to identify memory waste, underutilized executors, and inefficient cluster configurations.
Optimize EC2 instance types, node counts, EBS, and Spot/On-Demand fleet configurations, including R-family C-family migrations where appropriate.
Tune Spark/YARN configurations, including executor memory/cores, memory overhead, container sizing, dynamic allocation, and shuffle settings.
Perform end-to-end optimization experiments, deployments, testing, and metric validation using Amazon tools.
Monitor service health and ensure job performance, throughput, completion times, and SLAs are not degraded.
Collaborate with service teams and engineering leadership to present findings and implement optimization plans.
Create runbooks, SOPs, technical specifications, and reusable optimization playbooks for fleet-wide adoption.
Develop scalable frameworks to optimize hundreds of EMR clusters and enable other engineers to execute optimization workflows.