AWS Databricks Engineer
Location: Chicago, IL / Remote
Job type: Full time
Job Summary
We are seeking an experienced AWS Databricks Engineer to design, develop, and maintain scalable data engineering solutions using Databricks, Apache Spark, and AWS cloud services. The ideal candidate will have strong experience in data pipelines, ETL/ELT, data lake architecture, and cloud-based data processing.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using Databricks and Apache Spark.
- Build and optimize ETL/ELT workflows for batch and streaming data processing.
- Develop solutions using Databricks notebooks, Delta Lake, PySpark, and SQL.
- Implement and manage data solutions on AWS, including S3, Glue, Lambda, EMR, Redshift, and related services.
- Develop and maintain data lake/lakehouse architectures using Databricks and AWS.
- Perform data ingestion from databases, APIs, files, and other enterprise data sources.
- Implement Delta Lake features such as schema evolution, partitioning, optimization, and data versioning.
- Monitor, troubleshoot, and optimize data pipelines for performance, reliability, and scalability.
- Implement security, access controls, data governance, and best practices across AWS and Databricks environments.
- Work with data architects, analysts, developers, and business stakeholders to understand requirements and deliver data solutions.
- Implement CI/CD and deployment automation for Databricks and data engineering workloads.
- Develop unit/integration testing and ensure data quality and pipeline reliability.
Required Skills
- 5+ years of experience in Data Engineering.
- Strong hands-on experience with Databricks.
- Strong knowledge of Apache Spark and PySpark.
- Proficiency in Python and SQL.
- Strong experience with AWS cloud services, particularly:
- Amazon S3
- AWS Glue
- AWS Lambda
- Amazon Redshift
- Amazon EMR
- IAM
- Experience with Delta Lake and Lakehouse architecture.
- Strong understanding of ETL/ELT, data warehousing, and data lake concepts.
- Experience developing production-grade data pipelines.
- Experience with Git and CI/CD tools.
- Strong troubleshooting and performance-tuning skills.
Preferred Skills
- Experience with Databricks Workflows/Jobs and Unity Catalog.
- Experience with AWS Step Functions or Airflow.
- Experience with real-time/streaming technologies such as Kafka or Kinesis.
- Knowledge of Terraform or Infrastructure as Code.
- Experience with data governance, lineage, and security.
- Databricks or AWS certifications are a plus.
Education
Bachelor s degree in Computer Science, Information Technology, Engineering, or a related field preferred.