Hi,
Please find the job description below
Role: Data Bricks Engineer
Location: Plano, TX
Job Summary:
We are seeking an experienced Databricks Engineer with strong expertise in migration of PySpark/Hive workloads from AWS EMR or legacy platforms data platforms while ensuring data quality, security, and scalability
Key Responsibilities:
• Design, develop, and maintain scalable data pipelines using Databricks Implement and manage Unity Catalog for centralized data governanance
• Develop and optimize Spark Declarative Pipelines (SDP) and moderatePerform end-to-end data validation, reconciliation, and quality assurance.
• Lead migration projects involving:
• PySpark jobs from AWS EMR to Databricks Hive-based ETL workloads to Databricks and Delta Lake.
• Legacy data platforms to Databricks Lakehouse architecture.
• Convert and optimize Hive SQL, Spark SQL, and PySpark workloads forImple and automation for validation and PySparkjobs from AWS EMR to Databricks.
• Hive-based ETL workloads to Databricks and Delta Lake.
• Legacy data platforms to Databricks Lakehouse architecture.
• Convert and optimize Hive SQL, Spark SQL, and PySpark workloads for Database Implement data quality frameworks and automation for validation and reconcilation Work closely with architects, business stakeholders, and data consumers to | Optimize Databricks jobs for performance, reliability, and cost efficiency.
• Manage CI/CD processes and infrastructure automation for Databricks deplo Ensure compliance with enterprise data security, governance.
Required Skills:
• Databricks & Data Engineering Strong experience with Databricks Lakehouse Platform.
• Hands-on experience with Delta Lake, Medallion Architecture, and Databric Expertise in PySpark and Spark SQL.
• Experience working with large-scale data processing and performance tuning Unity Catalog Strong understanding of: Unity Catalog setup and administration.
• Strong understanding of Unity Catalog setup and administration Fine-grained access control (RBAC). Data lineage and audit capabilities.
• Data governance and security implementation.
• Spark Declarative Pipelines (SDP) Hands-on experience implementing and maintaining SDP pipelines.
• Knowledge of pipeline orchestration, monitoring, and optimization.
• Experience building reusable and scalable data transformation framework Data Validation & Quality Experience designing and implementing data validation frameworks.
• Data reconciliation between source and target systems. Validation of migrated datasets to ensure accuracy and completeness.
• Experience with data quality monitoring and observability tools.
• Migration Experience Proven experience in migratingPySparkworkloads from AWS EMR to Databricks.
• Mive-based ETL jobs to Databricks. Legacy data warehouses and Hadoop ecosystems to Databricks.