Amazon’s RFM team is seeking an engineer to optimize DRAM usage across EMR on EC2 clusters through cluster tuning, Spark/YARN optimization, and EC2 instance right-sizing. The role involves analyzing utilization, identifying over-provisioned resources, running optimization experiments, and delivering measurable memory/cost savings without impacting performance or SLAs.
Key Responsibilities
• Analyze EMR, Spark, YARN, and CloudWatch metrics to identify memory waste, underutilized executors, and inefficient cluster configurations.
• Optimize EC2 instance types, node counts, EBS, and Spot/On-Demand fleet configurations, including R-family → C-family migrations where appropriate.
• Tune Spark/YARN configurations, including executor memory/cores, memory overhead, container sizing, dynamic allocation, and shuffle settings.
• Perform end-to-end optimization experiments, deployments, testing, and metric validation using Amazon tools.
• Monitor service health and ensure job performance, throughput, completion times, and SLAs are not degraded.
• Collaborate with service teams and engineering leadership to present findings and implement optimization plans.
• Create runbooks, SOPs, technical specifications, and reusable optimization playbooks for fleet-wide adoption.
• Develop scalable frameworks to optimize hundreds of EMR clusters and enable other engineers to execute optimization workflows.