Sr. Databricks Engineer - Hybrid NYC

New York, NY, US • Posted 3 hours ago • Updated 31 minutes ago
Contract W2
On-site
USD70 - USD75/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Sr. Databricks Engineer - Hybrid NYC

Summary

job summary:

We are seeking an experienced Senior Data Engineer to design, develop, and optimize scalable data platforms that power enterprise analytics, machine learning, and AI-driven solutions. The ideal candidate will bring deep expertise in modern data engineering practices, cloud-native architectures, Lakehouse platforms, and distributed data processing technologies.



This role will play a critical part in building reliable, high-performance data ecosystems leveraging Databricks, Spark, Delta Lake, Snowflake, Kafka, and AWS, while also contributing to the adoption of Generative AI, Large Language Models (LLMs), Retrieval Augmented Generation (RAG), and AI-assisted engineering solutions.



Key Responsibilities



Data Platform Engineering



Design, build, and maintain scalable batch and real-time data pipelines.



Develop and optimize data ingestion, transformation, and processing frameworks for structured, semi-structured, and unstructured datasets.



Implement modern Lakehouse architectures utilizing Databricks, Delta Lake, and Medallion (Bronze, Silver, Gold) design patterns.



Build data solutions that support enterprise analytics, reporting, and machine learning initiatives.



Ensure data quality, governance, lineage, security, and compliance across data ecosystems.



Big Data & Streaming Solutions



Develop distributed data processing applications using PySpark and Spark SQL.



Build and maintain streaming pipelines using Kafka, Spark Structured Streaming, and AWS Kinesis.



Design fault-tolerant, scalable systems capable of processing large data volumes with low latency.



Optimize workload performance through partitioning strategies, clustering, caching, and query tuning.



Cloud & Lakehouse Architecture



Architect and implement cloud-based data solutions on AWS.



Utilize AWS services including S3, EMR, EC2, Athena, Redshift, RDS, Lambda, IAM, SNS, and SQS.



Design data storage and processing strategies that maximize reliability while minimizing operational costs.



Support migration initiatives from traditional Hadoop and EMR environments to modern cloud-native platforms.



Data Operations & Automation



Develop orchestration and scheduling frameworks using Airflow and Databricks Workflows.



Build CI/CD pipelines and automation frameworks for deployment, monitoring, and data platform operations.



Collaborate closely with architects, analysts, data scientists, and business stakeholders to deliver enterprise-grade solutions.



AI & Intelligent Platform Engineering



Implement Generative AI-powered solutions for engineering productivity and operational excellence.



Develop applications leveraging Large Language Models (LLMs), Retrieval Augmented Generation (RAG), Vector Databases, and Model Context Protocol (MCP).



Build AI-assisted documentation, developer productivity tooling, and intelligent platform capabilities.



Evaluate emerging AI technologies and identify opportunities for adoption within data engineering processes.



Required Qualifications



Bachelor's or Master's degree in Computer Engineering, Computer Science, Information Systems, or a related field.



8+ years of experience in software engineering, data engineering, or big data platform development.



Strong experience designing and implementing enterprise-scale data pipelines.



Hands-on expertise with:



Python



PySpark



Spark SQL



SQL



Kafka



Databricks



Delta Lake



Snowflake



Hive



Experience building data solutions on AWS cloud platforms.



Strong understanding of distributed computing, data modeling, and large-scale data processing.



Experience with Git-based development workflows and CI/CD practices.



Excellent analytical, troubleshooting, and problem-solving skills.



Preferred Qualifications



Experience with real-time streaming architectures and event-driven systems.



Knowledge of data governance, metadata management, and data quality frameworks.



Experience with generative AI technologies including:



LLMs



RAG



Vector Databases



AI Agents



MCP integrations



Experience developing developer productivity tools and AI-assisted engineering workflows.



Exposure to enterprise supply chain, retail, healthcare, or manufacturing data domains.



AWS certifications are highly preferred.



Technical Skills



Programming Languages



Python



Java



SQL



Shell Scripting



C/C++



Big Data & Data Engineering



PySpark



Spark SQL



Hive



Databricks



Delta Lake



Snowflake



Kafka



HBase



Sqoop



Workflow & Orchestration



Apache Airflow



Databricks Workflows



Oozie



Cloud Technologies



AWS S3



EMR



EC2



Athena



Redshift



RDS



IAM



Lambda



SNS



SQS



AI & Modern Engineering



Generative AI



Large Language Models (LLMs)



Retrieval Augmented Generation (RAG)



Agentic AI Systems



Model Context Protocol (MCP)



Vector Databases



Visualization & Tools



Tableau



Git



Docker



Splunk



IntelliJ IDEA



PyCharm



Cursor



Preferred Certifications



AWS Certified Solutions Architect - Associate



AWS Certified Cloud Practitioner



Databricks Certifications (preferred)



What Success Looks Like



Deliver highly scalable and reliable data pipelines.



Improve platform performance, efficiency, and cost optimization.



Enable enterprise-wide analytics and AI initiatives through trusted data products.



Drive modernization of data platforms and adoption of cloud-native architectures.



Leverage AI technologies to enhance engineering efficiency, automation, and innovation.



Ideal Candidate Profile: A senior-level data engineer with extensive experience in Databricks, Spark, AWS, Kafka, Snowflake, and Lakehouse architectures, who is equally passionate about modern AI technologies and building intelligent data platforms for the future.







location: New York, New York

job type: Contract

salary: $70 - 75 per hour

work hours: 9am to 6pm

education: Bachelors



responsibilities:

We are seeking an experienced Senior Data Engineer to design, develop, and optimize scalable data platforms that power enterprise analytics, machine learning, and AI-driven solutions. The ideal candidate will bring deep expertise in modern data engineering practices, cloud-native architectures, Lakehouse platforms, and distributed data processing technologies.



This role will play a critical part in building reliable, high-performance data ecosystems leveraging Databricks, Spark, Delta Lake, Snowflake, Kafka, and AWS , while also contributing to the adoption of Generative AI, Large Language Models (LLMs), Retrieval Augmented Generation (RAG), and AI-assisted engineering solutions .



Key Responsibilities



Data Platform Engineering




  • Design, build, and maintain scalable batch and real-time data pipelines.


  • Develop and optimize data ingestion, transformation, and processing frameworks for structured, semi-structured, and unstructured datasets.


  • Implement modern Lakehouse architectures utilizing Databricks, Delta Lake, and Medallion (Bronze, Silver, Gold) design patterns.


  • Build data solutions that support enterprise analytics, reporting, and machine learning initiatives.


  • Ensure data quality, governance, lineage, security, and compliance across data ecosystems.



Big Data & Streaming Solutions






  • Develop distributed data processing applications using PySpark and Spark SQL.


  • Build and maintain streaming pipelines using Kafka, Spark Structured Streaming, and AWS Kinesis.


  • Design fault-tolerant, scalable systems capable of processing large data volumes with low latency.


  • Optimize workload performance through partitioning strategies, clustering, caching, and query tuning.



Cloud & Lakehouse Architecture






  • Architect and implement cloud-based data solutions on AWS.


  • Utilize AWS services including S3, EMR, EC2, Athena, Redshift, RDS, Lambda, IAM, SNS, and SQS.


  • Design data storage and processing strategies that maximize reliability while minimizing operational costs.


  • Support migration initiatives from traditional Hadoop and EMR environments to modern cloud-native platforms.



Data Operations & Automation






  • Develop orchestration and


Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: cxsapwma1
  • Position Id: 1346842
  • Posted 3 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

New York, New York

Today

Easy Apply

Full-time, Part-time, Third Party, Contract

USD 65-65

Hybrid in New York, New York

Today

Easy Apply

Contract

Depends on Experience

Hybrid in New York, New York

16d ago

Easy Apply

Contract

$65 - $75

New York, New York

7d ago

Easy Apply

Full-time

60 - 70

Search all similar jobs