Location - NYC - 3 days onsite
Azure Data lead - Python, Pyspark, Databricks, ADF, Data Lake,
Azure Senior Data Lead Leads modernization of Python applications into scalable PySpark solutions on Azure Databricks Role Purpose Lead the modernization and migration of existing Python object-oriented applications into scalable PySpark and Spark SQL data-processing solutions on Azure Databricks bringing a strong blend of software engineering, data engineering, cloud architecture, and performance optimization. Key Responsibilities Analyze existing Python OOP applications and redesign single-node processing logic for distributed Spark execution. Design, develop, and deploy enterprise-scale data pipelines on Azure Databricks; build reusable PySpark frameworks and utility modules. Implement Delta Lake solutions using the Bronze Silver Gold architecture. Build robust ETL/ELT pipelines with Azure Data Factory, ADLS Gen2, and Azure Synapse Analytics. Implement data quality, reconciliation, validation, and monitoring frameworks. Optimize Spark jobs (partitioning, bucketing, caching, broadcast joins, Adaptive Query Execution, Delta optimization) and benchmark converted applications against original Python implementations. Core Skills Python (expert), OOP, and advanced Python design patterns PySpark, Spark SQL, and SQL Azure Databricks, Azure Data Factory, ADLS Gen2 Apache Spark, Delta Lake, Data Lakehouse architecture, distributed computing Must-have (per requisition): Python, Azure Databricks, Azure Data Factory (ADF), MS SQL, Oracle PL/SQL. Good to have: PySpark; certifications in Azure Data Factory, Azure Databricks, SQL, Oracle, or Python. Experience & Expected Outcome Senior data engineering leader with proven delivery of large-scale Databricks modernization programs. Expected outcome: existing Python applications converted into scalable, cost-efficient, enterprise-grade data solutions on Azure Databricks with proven performance parity.