Job Description
Position: Senior Data Engineer / Data Platform Engineer
We are looking for a Senior Data Engineer with strong experience building and supporting large-scale data platforms, distributed data pipelines, and real-time data processing systems. The ideal candidate should be hands-on with Spark, Kafka, Databricks, Airflow, and cloud data technologies and comfortable working with high-volume enterprise data environments.
Key Responsibilities
Design, develop, and maintain scalable batch and real-time data pipelines.
Build distributed data processing solutions using Apache Spark and Spark Structured Streaming.
Develop streaming and event-driven data pipelines using Apache Kafka/Kafka Connect.
Build and optimize ETL/ELT workflows using Python, Java, and SQL.
Work with Databricks and modern lakehouse technologies such as Delta Lake or Apache Hudi.
Develop and manage data workflows using Apache Airflow.
Design data models, partitioning strategies, schema evolution, and data quality frameworks.
Monitor pipeline performance, data freshness, reliability, and operational health.
Work with cloud platforms such as Google Cloud Platform or Azure and technologies including BigQuery, Cloud Storage, Dataflow, Kubernetes, or similar.
Collaborate with analytics, ML, product, and infrastructure teams to build scalable data solutions.
Participate in architecture discussions, code reviews, troubleshooting, and performance optimization.
Implement CI/CD, automated testing, and engineering best practices for data platforms.
Required Skills
8+ years of experience in Data Engineering or related roles.
Strong hands-on experience with Apache Spark / Spark Structured Streaming.
Strong experience with Kafka and real-time/event-driven data pipelines.
Strong experience with Databricks.
Strong experience with Python and SQL; Java or Scala is highly preferred.
Experience with Apache Airflow or similar workflow orchestration tools.
Experience building large-scale ETL/ELT pipelines.
Experience with cloud data platforms, preferably Google Cloud Platform or Azure.
Experience with data lakes/lakehouse architectures.
Strong understanding of distributed systems, data modeling, performance tuning, and data quality.
Experience with Git and CI/CD tools such as Jenkins, GitHub Actions, or similar.
Preferred
Experience with Delta Lake or Apache Hudi.
Experience with BigQuery, GCS, Dataflow, Cassandra, Presto/Trino, or similar technologies.
Experience with Kubernetes and Docker.
Experience with streaming data and CDC.
Experience working with ML/analytics data platforms.
Experience establishing data contracts, SLAs, monitoring, and observability.
Enterprise-scale data platform experience with very large datasets.
Strong communication and stakeholder-management skills.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: wesca004
- Position Id: JOB-7388
- Posted 14 hours ago