Title: Azure Data Lead
Location: NYC (3 days onsite)
Skills: Azure Data lead - Python, Pyspark, Databricks, ADF, Data Lake
Introduction
The Azure Data Lead will be responsible for leading the modernization and migration of existing Python object-oriented applications into scalable PySpark and Spark SQL data-processing solutions on Azure Databricks. This role requires a strong blend of software engineering, data engineering, cloud architecture, and performance optimization.
Responsibilities
- Analyze existing Python OOP applications and redesign single-node processing logic for distributed Spark execution.
- Design, develop, and deploy enterprise-scale data pipelines on Azure Databricks; build reusable PySpark frameworks and utility modules.
- Implement Delta Lake solutions using the Bronze–Silver–Gold architecture.
- Build robust ETL/ELT pipelines with Azure Data Factory, ADLS Gen2, and Azure Synapse Analytics.
- Implement data quality, reconciliation, validation, and monitoring frameworks.
- Optimize Spark jobs (partitioning, bucketing, caching, broadcast joins, Adaptive Query Execution, Delta optimization) and benchmark converted applications against original Python implementations.
Requirements
Core Skills:
- Python (expert), OOP, and advanced Python design patterns
- PySpark, Spark SQL, and SQL
- Azure Databricks, Azure Data Factory, ADLS Gen2
- Apache Spark, Delta Lake, Data Lakehouse architecture, distributed computing
Must-have: Python, Azure Databricks, Azure Data Factory (ADF), MS SQL, Oracle PL/SQL.
Good to have: PySpark; certifications in Azure Data Factory, Azure Databricks, SQL, Oracle, or Python.
Experience & Expected Outcome
The ideal candidate for this role will be a senior data engineering leader with proven delivery of large-scale Databricks modernization programs. The expected outcome is to have existing Python applications converted into scalable, cost-efficient, enterprise-grade data solutions on Azure Databricks with proven performance parity.