Job Description
We are seeking a highly skilled Senior PySpark Developer to support a large-scale cloud transformation initiative focused on building modern data platforms on Cloudera, Apache Iceberg, Airflow, and AWS. The role will be responsible for designing, developing, and optimizing scalable data pipelines that ingest, transform, and process large volumes of enterprise data.
The successful candidate will work closely with data architects, business analysts, cloud engineers, and application teams to build reliable, high-performance data solutions supporting analytics, reporting, and operational workloads. Responsibilities include developing PySpark applications, implementing ETL/ELT processes, designing Iceberg-based data models, orchestrating workflows using Airflow, and leveraging AWS services to support cloud-native data processing.
This requires strong expertise in distributed data processing, performance tuning, data quality, and cloud-based architectures. The candidate will participate in end-to-end solution delivery, from requirements analysis and solution design through development, testing, deployment, and production support. They will also contribute to establishing best practices, reusable frameworks, CI/CD processes, and data governance standards across the platform.
The ideal candidate is a hands-on developer with deep technical expertise in modern data engineering technologies and experience working in large-scale enterprise cloud migration and modernization programs.
Required Qualifications
- Bachelor''s degree in Computer Science, Information Systems, Engineering, or related field.
- 10+ years of IT experience with 6+ years in Data Engineering and Big Data development.
- Strong hands-on experience with PySpark and Spark-based data processing.
- Experience developing data pipelines on Cloudera Data Platform (CDP).
- Strong knowledge of Apache Iceberg table architecture and data lake concepts.
- Experience designing and managing workflows using Apache Airflow.
- Experience with SQL and data modeling techniques.
- Strong understanding of ETL/ELT development and data transformation frameworks.
- Experience working with large-scale structured and unstructured datasets.
- Knowledge of Git, CI/CD pipelines, and DevOps practices.
- Strong analytical, troubleshooting, and problem-solving skills.
- Excellent communication and collaboration skills.
Preferred Qualifications
- Experience with AWS services including S3, Glue, Lambda, ECS/EKS, EMR, RDS, IAM, and CloudWatch.
- Experience building cloud-native data lake and Lakehouse solutions.
- Strong knowledge of data partitioning, performance tuning, and query optimization.
- Experience with Kafka or event-driven data processing frameworks.
- Experience with Java or Python application development.
- Experience implementing data quality, metadata, and data governance solutions.
- Healthcare, Medicaid, Claims Processing, or Insurance industry experience.
- Experience working in Agile development environments.
- AWS Certification and/or Cloudera Certification preferred.
- Experience supporting large-scale cloud migration and modernization initiatives.