**** Initial webcam interview followed by In-Person interview *** Long term project with possible contract to hire ***Linkedin Must*** Onsite ****
Job summary:
We are seeking a highly skilled Data Engineer to join a team responsible for building and enhancing a large-scale decision platform that drives customer-focused business decisions. The ideal candidate will have strong expertise in Spark, Java, AWS, and large-scale data processing environments.
Key responsibilities:
- Design, develop, and optimize scalable data pipelines for ingesting and processing high-volume datasets.
- Build and maintain distributed data processing solutions using Apache Spark and Java.
- Process and manage 20M+ to 40M+ daily data records efficiently.
- Develop batch and real-time data processing workflows.
- Work with event-driven architectures and streaming platforms.
- Collaborate with cross-functional teams to enhance data platform capabilities.
- Implement and manage cloud-native solutions within AWS environments.
- Utilize Infrastructure as Code (IaC) methodologies using Terraform.
Required skills:
- 5 to 10 years of experience.
- Strong experience with Apache Spark (Must Have)
- Strong experience with Java (Must Have)
- Hands-on experience with AWS Cloud Services (Must Have)
- Experience with Terraform for Infrastructure as Code (Must Have)
- Experience building data ingestion and data transformation pipelines
- Experience working with large-scale distributed data environments
- Knowledge of real-time and batch data processing architectures
Preferred skills:
Cloud & Technology Stack:
- Apache Spark
- Java
- AWS EKS
- AWS EMR
- AWS S3
- AWS Aurora
- AWS MSK
- AWS SNS
- AWS SQS
- Apache Kafka
- Terraform
Ideal candidate:
- Strong background in data engineering and distributed computing.
- Experience handling high-volume data platforms.
- Comfortable working in cloud-native, event-driven architectures.
- Excellent problem-solving and analytical skills.
Must have:
- Pyspark
- Snowflake
- Databricks
- AWS