Must Have Skills
AWS Glue, Lambda, Redshift, EMR
Nice to have skills
SQL, Python/Pyspark, DAG
Detailed Job Description
Architect and implement cloud native data pipelines on AWS.
Build serverless compute layers using Lambda, Step Functions, and Airflow.
Design data lake storage strategies using S3, Glue, Athena, and Redshift.
Integrate real-time data streams via Kinesis and SQS
Collaborate with business analysts to translate reporting requirements into schema and pipeline design
Own data quality, pipeline monitoring, and alerting via CloudWatch
Lead schema design for Athena and Postgres tables DDL, partitioning, indexing
Evaluate and integrate AI/ML services Bedrock, Comprehend, Transcribe
Document architecture decisions and maintain runbooks
5 years AWS data engineering experience Expert Python data transformation, boto3, async processing Deep knowledge of S3, Lambda, SQS, Kinesis, Step Functions, EventBridge, DynamoDB, Glue, Athena, Redshift, Airflow
Airflow production DAG development and debugging SQL complex queries, window functions, partitioning, performance tuning IAM roles, policies, cross account access, permissions boundaries
Minimum years of experience
5-8 years