Pipeline Engineering: Develop and maintain scalable ELT pipelines using reusable ingestion frameworks that support both batch and event-driven processing across multiple data sources such as ERPs, APIs, vendor feeds, and relational systems.
AWS Data Platform Development: Design and enhance AWS-native data solutions aligned with a Medallion (Bronze/Silver/Gold) Lakehouse architecture, with an emphasis on performance optimization and cost management.
Infrastructure & DevOps: Build and manage AWS infrastructure using Terraform, and support CI/CD processes through GitOps methodologies, including automated testing, monitoring, alerting, and system recovery capabilities.
Data Modeling & Transformation: Create and maintain dimensional models and Gold-layer datasets using SQL, Python, and PySpark. Implement scalable ingestion processes with strong handling of schema evolution, auditing, and performance tuning techniques such as partitioning, clustering, and materialization.
Data Quality & Governance: Integrate automated data validation, anomaly detection, and lineage tracking into pipelines, while contributing to metadata management practices.
Reporting & BI Enablement: Diagnose and resolve complex data issues, ensuring pipelines deliver data optimized for Power BI. Work closely with BI developers to align data models with reporting needs and troubleshoot dashboard-related challenges.
Cross-Functional Collaboration: Partner with architecture and analytics teams to translate business requirements into technical solutions, and actively participate in code reviews, sprint planning, and architectural discussions.
Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related discipline, or equivalent practical experience.
At least 7 years of experience in data engineering, with strong hands-on expertise in AWS services and distributed data technologies.
Advanced proficiency in Python and SQL, including experience with Spark/PySpark.
Proven experience building and maintaining AWS-based data pipelines using services such as Glue, Step Functions, Lambda, S3, Athena, SNS, SQS, and Redshift.
Experience with event-driven data architectures.
Hands-on experience processing and managing large-scale vendor data feeds.
Practical knowledge of Medallion architecture within a data lake or Lakehouse environment.
Experience using Terraform for infrastructure-as-code deployments.
Familiarity with CI/CD tools such as Bitbucket, GitHub, or AWS CodePipeline for pipeline deployment.
Strong understanding of data warehousing principles, including star schema design, dimensional modeling, and slowly changing dimensions (SCD).
Experience integrating data with Power BI or comparable business intelligence tools.
Working knowledge of UNIX/Linux environments, including shell scripting.
Experience supporting production systems, including monitoring, troubleshooting, and on-call responsibilities.
Familiarity with Agile development methodologies.
Experience integrating data from Oracle EBS.
Familiarity with data quality tools such as Great Expectations or dbt testing frameworks.
Experience with AWS CDK or CloudFormation alongside Terraform.
Knowledge of data cataloging and lineage tools such as Alation or Collibra.
AWS certifications (e.g., AWS Certified Data Engineer, AWS Certified Solutions Architect).