DataBrick Data Engineer

Remote • Posted 8 hours ago • Updated 8 hours ago
Full Time
Remote
Fitment

Dice Job Match Score™

🛠️ Calibrating flux capacitors...

Job Details

Skills

  • Extract
  • Transform
  • Load
  • Legacy Systems
  • Code Refactoring
  • Optimization
  • Migration
  • Apache Spark
  • Caching
  • Unity
  • Datastage
  • IBM InfoSphere DataStage
  • PySpark
  • Python
  • SQL
  • Shell Scripting
  • Cloud Computing
  • DevOps
  • Amazon Web Services
  • Continuous Integration
  • Continuous Delivery
  • Orchestration
  • Apache Airflow
  • Databricks
  • Workflow

Summary

Job Title: DataBricks Data Engineer

Remote

Full Time Only

Job Summary

The Databricks Data Engineer owns the end-to-end migration strategy, target architecture design, and technical execution of moving legacy ETL workloads to the Databricks Lakehouse. He will establish migration standards, optimize PySpark pipelines, orchestrate complex data workflows, and deploy proprietary automation tools to ensure a seamless, high-performing transition from legacy systems.

Key Responsibilities

Architectural Strategy & Governance

  • Design the target Databricks Lakehouse architecture utilizing Delta Lake, Photon, and Unity Catalog.
  • Establish global code refactoring standards, optimization benchmarks, and PySpark best practices.
  • Resolve highly complex dependency mappings and architect seamless, zero-downtime dual-run strategies.
  • Lead the technical deployment and integration of specialized migration accelerators

Hands-on Engineering & Optimization

  • Review automated output from migration tools and manually refactor complex legacy logic into high-performing PySpark notebooks.
  • Eliminate legacy anti-patterns such as massive row-by-row processing and inefficient lookups.
  • Optimize PySpark code performance using advanced Spark features including Z-Ordering, partitioning, and caching.
  • Build robust Databricks Workflows and orchestrate complex DAGs based on comprehensive source lineage.

Technical Skills & Competencies

  • Core Platforms: Databricks, Delta Lake, Unity Catalog, Photon, DataStage.
  • Languages & Frameworks: PySpark, Python, SQL, Shell Scripting.
  • Cloud & DevOps: AWS alongside CI/CD deployment pipelines.
  • Orchestration: Apache Airflow, Databricks Workflows.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10236892
  • Position Id: OOJ - 5532-4533-1785450688
  • Posted 8 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Full-time

Remote

20d ago

Easy Apply

Contract

Depends on Experience

Remote

2d ago

Easy Apply

Full-time

65 - 75

Remote

3d ago

Easy Apply

Contract

50

Search all similar jobs