Design, develop, and maintain scalable ETL/ELT data pipelines using Python, PySpark, Apache Spark, and Databricks.
Develop and optimize data processing workflows using Azure Databricks and Azure Data Factory (ADF).
Write complex and optimized SQL queries for data extraction, transformation, validation, and analysis.
Build batch and large-scale data processing solutions using Apache Spark/PySpark.
Develop reusable data engineering frameworks and components using Python.
Create and maintain ADF pipelines, datasets, triggers, and integrations.
Perform data transformation, cleansing, validation, and reconciliation across multiple data sources.
Work with Unix/Linux shell scripting for automation, scheduling, and production support.
Use Control-M for enterprise job scheduling, monitoring, and batch processing.
Troubleshoot production data pipelines and resolve performance, data quality, and integration issues.
Optimize Spark jobs, Databricks workloads, SQL queries, and ETL processes for performance and scalability.
Collaborate with architects, analysts, developers, and business teams to understand data requirements.
Follow data engineering best practices for performance, reliability, scalability, and data quality.
Support deployment, monitoring, and production operations of data pipelines.