Job Summary:
We are seeking a skilled, experienced and motivated Data Engineer with expertise in AWS services such as S3, Redshift, Glue, Step Functions, and Lambda to join our support team. The ideal candidate will be responsible for managing, monitoring, and troubleshooting data pipelines and workflows, ensuring the seamless operation of data infrastructure, and providing ongoing support to maintain data availability and integrity.
Minimum 7 yrs in Data engineering with relevant skills
Mandatory Skills:
AWS Services (S3, Redshift, Lambda, Glue, Step function) Python
Technical Expertise:
Proficient in AWS services: S3, Redshift, Glue, Step Functions, Lambda.
Strong understanding of ETL/ELT processes and data transformation.
Experience with monitoring and debugging data pipelines in a production environment.
Database Knowledge:
Proficiency in SQL and hands-on experience with Redshift for data modeling and performance tuning.
Knowledge of other databases (e.g., PostgreSQL, MySQL) is a plus.
Programming Skills:
Proficiency in Python for scripting, data manipulation, and building serverless applications with Lambda.
Knowledge of PySpark or similar frameworks is a plus.
Key Responsibilities:
AWS Data Pipeline Maintenance: Monitor and troubleshoot data pipelines and workflows utilizing AWS Glue, Step Functions, and Lambda. Optimize existing data pipelines for performance and cost efficiency.
Data Storage Management: Manage and support data stored in Amazon S3 and ensure efficient storage policies (e.g., lifecycle rules, versioning, and encryption). Ensure optimal performance and availability of Amazon Redshift clusters, including schema maintenance and query tuning.
Issue Resolution: Respond to incidents related to data failures, latency, or data quality and provide timely resolutions. Debug and resolve issues in ETL jobs, data ingestion, and transformations.
Performance Optimization: Perform root cause analysis for recurring issues and implement solutions to enhance the reliability of the data ecosystem. Identify areas for process improvement and implement automation wherever feasible.
Data Quality and Governance: Implement monitoring tools and dashboards to track data pipeline health. Collaborate with stakeholders to ensure adherence to data security, governance, and compliance policies.
Documentation and Reporting: Maintain up-to-date documentation for data pipelines, workflows, and troubleshooting steps. Provide regular reports on system performance and key metrics.
Collaboration and Support: Work closely with data engineering, analytics, and operations teams to resolve issues and gather requirements for enhancements. Support ad-hoc data requests and ensure timely delivery of data to business users.
Qualification : Any Graduation