Position: Azure Data lead
Location: New York City , NY (Hybrid)
Must-have (per requisition): Python, Azure Databricks, Azure Data Factory (ADF), MS SQL, Oracle PL/SQL
Azure Data lead - Python, Pyspark, Databricks, ADF, Data Lake
Azure Senior Data Lead Leads modernization of Python applications into scalable PySpark solutions on Azure Databricks
Role Purpose: Lead the modernization and migration of existing Python object-oriented applications into scalable PySpark and Spark SQL data-processing solutions on Azure Databricks — bringing a strong blend of software engineering, data engineering, cloud architecture, and performance optimization.
Key Responsibilities
• Analyze existing Python OOP applications and redesign single-node processing logic for distributed Spark execution.
• Design, develop, and deploy enterprise-scale data pipelines on Azure Databricks; build reusable PySpark frameworks and utility modules.
• Implement Delta Lake solutions using the Bronze–Silver–Gold architecture.
• Build robust ETL/ELT pipelines with Azure Data Factory, ADLS Gen2, and Azure Synapse Analytics.
• Implement data quality, reconciliation, validation, and monitoring frameworks.
• Optimize Spark jobs (partitioning, bucketing, caching, broadcast joins, Adaptive Query Execution, Delta optimization) and benchmark converted applications against original Python implementations.
Core Skills
• Python (expert), OOP, and advanced Python design patterns
• PySpark, Spark SQL, and SQL
• Azure Databricks, Azure Data Factory, ADLS Gen2
• Apache Spark, Delta Lake, Data Lakehouse architecture, distributed computing
Good to have: PySpark; certifications in Azure Data Factory, Azure Databricks, SQL, Oracle, or Python.
Experience & Expected Outcome Senior data engineering leader with proven delivery of large-scale Databricks modernization programs.
Expected outcome: existing Python applications converted into scalable, cost-efficient, enterprise-grade data solutions on Azure Databricks with proven performance parity.