Data Engineer Role

Remote • Posted 6 hours ago • Updated 6 hours ago
Full Time
No Travel Required
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • python
  • databricks
  • pyspark

Summary

Role- Data Engineer

Location- New York

 

Job Description

We are seeking an experienced Data Engineer to join our team and build robust, scalable data pipelines. In this role, you will:

• Design and implement scalable PySpark data pipelines for batch and streaming workloads

• Optimize Spark jobs and queries for performance and cost efficiency

• Build and maintain ETL/ELT processes following data engineering best practices

• Troubleshoot and resolve complex data pipeline and processing issues

• Collaborate with data teams to ensure data quality and reliability

Top Skills

Databricks Platform Experience

Data Engineering & Pipeline Development

• Advanced ETL/ELT pipeline design and development

• Incremental data processing patterns (CDC, SCD Type 2)

Data Processing & Optimization

• Spark optimization techniques (partitioning, bucketing, caching, broadcast joins)

Required Technical Skills

Databricks & Spark Proficiency: 3+ years of hands-on experience building data pipelines in Databricks; deep understanding of Spark fundamentals, transformations, actions, and performance optimization techniques including partitioning, caching, and resource management

Advanced PySpark and SQL Skills: Expert-level proficiency writing production-quality PySpark code and complex SQL queries for data transformation, aggregation, and analysis; experience with DataFrame API, Spark SQL, and UDFs; strong understanding of lazy evaluation and execution plans

Data Engineering & ETL/ELT: Proven experience building and maintaining production data pipelines; hands-on experience with incremental data loading, change data capture (CDC), and slowly changing dimensions; experience handling data quality issues and implementing data validation frameworks

Cloud & Big Data Technologies: Strong proficiency with AWS services (S3, EC2, IAM, Glue, Athena); experience working with large-scale distributed data processing; familiarity with data formats (Parquet, Delta, JSON, Avro) and compression techniques

DevOps & CI/CD: Experience with version control (Git) and CI/CD pipelines using GitLab, GitHub Actions, or similar tools; familiarity with testing data pipelines and deployment automation; experience with Databricks Repos and workspace-level integrations

Data Governance: Understanding of data lineage, cataloging, and metadata management; experience implementing data quality checks and monitoring; knowledge of data privacy and security best practices in cloud environments (nice to have)

If interested please email updated resume to

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10115720
  • Position Id: 9089483
  • Posted 6 hours ago
Contact the job poster
DS

Dipa Shetty

Recruiter @ Quinnox Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote or Texas

Today

Easy Apply

Full-time, Contract, Third Party

Remote or Pennsylvania

Today

Full-time

USD 20.00 per hour

Remote

9d ago

Easy Apply

Full-time

130,000 - 150,000

Remote

2d ago

Easy Apply

Full-time, Third Party

$70 - $80

Search all similar jobs