Note:
This role requires you to travel across USA (Ideally 15 days/month)
About the Role:
We are seeking an experienced Databricks Architect to design, develop, and optimize scalable data engineering solutions on the Databricks Lakehouse platform. This role focuses on leveraging hands-on expertise in Unity Catalog, Apache Spark, PySpark, SQL, and Delta Lake to build robust data pipelines while ensuring industry-leading data governance, security, and access control standards across cloud environments.
Key Responsibilities:
- Design and develop scalable data pipelines using Databricks, Apache Spark, PySpark, and SQL.
- Build and optimize ETL/ELT pipelines using Delta Lake and Delta Tables.
- Implement and manage Unity Catalog across Databricks environments.
- Configure and manage Catalogs, Schemas, Tables, Views, External Locations, and Storage Credentials.
- Implement data access controls, RBAC, permissions, and governance policies using Unity Catalog.
- Support data lineage, auditing, data discovery, and governance requirements.
- Design and maintain secure data lakehouse architectures across cloud environments.
- Develop Databricks Jobs, Workflows, notebooks, and production-grade data pipelines.
- Optimize Spark jobs, SQL queries, Delta tables, and overall pipeline performance.
- Implement data quality, validation, monitoring, and error-handling frameworks.
- Work with cloud storage platforms such as Azure Data Lake Storage, Amazon S3, or Google Cloud Storage.
- Integrate Databricks with enterprise data platforms, databases, APIs, and downstream applications.
- Implement CI/CD processes for Databricks development and deployment.
- Use Git, Terraform, or Databricks Asset Bundles for infrastructure and deployment automation.
- Troubleshoot production data pipelines and resolve performance and data-quality issues.
- Collaborate with Data Architects, Cloud Engineers, Data Scientists, Business Analysts, and application teams.
Required Qualifications:
- 5+ years of experience in Data Engineering.
- Strong hands-on experience with Databricks.
- Strong experience implementing and working with Unity Catalog.
- Advanced PySpark and Apache Spark skills.
- Strong SQL development and query optimization experience.
- Hands-on experience with Delta Lake / Delta Tables.
- Experience building enterprise-grade ETL/ELT data pipelines.
- Experience with data governance, security, access management, and data lineage.
- Experience with at least one major cloud platform: Azure, AWS, or Google Cloud Platform.
- Strong understanding of Lakehouse architecture and modern data platforms.
- Experience with Databricks Workflows/Jobs and production deployments.
- Experience with Git and CI/CD.
- Good understanding of data modeling and performance optimization.
Preferred Qualifications:
- Experience with Terraform or Databricks Asset Bundles.
- Experience with Azure Data Lake Storage, Amazon S3, or Google Cloud Storage.
- Experience with Azure Data Factory, AWS Glue, Airflow, or similar orchestration tools.
- Experience implementing Unity Catalog migration from legacy Hive Metastore.
- Experience with row-level and column-level security.
- Experience with Dynamic Views and fine-grained data access.
- Experience with Databricks SQL and SQL Warehouses.
- Experience with Delta Live Tables / Lakeflow Declarative Pipelines.
- Experience with CI/CD automation for Databricks.