Role Overview
We are seeking a highly technical Data Architect with strong hands-on experience in Azure Databricks, Data Modeling, Spark/PySpark, SQL, Delta Lake, and modern data architecture. The role focuses on designing scalable data architecture while reviewing and troubleshooting data engineering implementations.
Key Responsibilities
- Design scalable Lakehouse and Medallion Architecture using Azure Databricks.
- Design Conceptual, Logical, and Physical Data Models for MedTech Supply Chain data.
- Define architecture for ingestion, transformation, curation, and downstream consumption.
- Design batch, incremental, CDC, and streaming data-processing patterns.
- Develop and review solutions using PySpark, Spark SQL, and advanced SQL.
- Design reusable and metadata-driven data engineering frameworks.
- Review source-to-target mappings and complex transformation logic.
- Define data-quality, reconciliation, lineage, metadata, and audit requirements.
- Design and review Delta Lake implementations including MERGE, schema evolution, and historical processing.
- Review Spark workloads for performance, scalability, and cost optimization.
- Define security and governance using Unity Catalog.
- Establish Databricks CI/CD and deployment standards.
- Work closely with Data Engineers, Data Modelers, Product Managers, and business teams.
Core Technical Skills
- Azure Databricks
- Apache Spark
- PySpark
- Advanced SQL / Spark SQL
- Delta Lake
- Unity Catalog
- Databricks Workflows / Jobs
- Databricks Asset Bundles
- Lakeflow / DLT
- Auto Loader
- CDC / Incremental Processing
- Structured Streaming
- Delta MERGE
- Schema Evolution
- Medallion Architecture
- Data Modeling
- Data Quality
- Metadata & Data Lineage
- Performance Tuning
Data Modeling & Architecture
Strong experience with Conceptual, Logical, and Physical Data Modeling; Dimensional Modeling; Fact and Dimension Design; Star / Snowflake Schema; Data Vault; SCD Type 1 / Type 2; Master and Reference Data; Enterprise and Domain Data Models.
Modern Data & AI Architecture
Understanding of Vector Databases and Vector Search, embeddings, semantic search, RAG architecture, Knowledge Graphs, ontology and semantic modeling, and integration of structured/unstructured data with Lakehouse, Vector Search, Knowledge Graph, and AI applications.
MedTech Supply Chain Knowledge
- Materials
- Plants & Storage Locations
- Inventory
- Sales Orders
- Purchase Orders
- Deliveries
- Shipments
- Material Movements
- Manufacturing
- Suppliers & Customers
- Demand & Supply Planning
- ATP
- Lead Times
- Order Fulfillment
- Strong SAP Supply Chain data knowledge is preferred.
Required Experience
- 10+ years of Data Engineering / Data Architecture experience.
- Strong hands-on Azure Databricks, PySpark, SQL, Data Modeling, and Lakehouse Architecture experience.
- Experience designing large-scale enterprise data platforms.
- Experience with Delta Lake, Unity Catalog, metadata-driven pipelines, and data-quality frameworks.
- Good understanding of Vector Databases, Knowledge Graphs, and Ontology.