Senior Data Engineer – Cloud, Big Data & Data Platforms


GTSS Inc
Dice Job Match Score™
🔢 Crunching numbers...
Job Details
Skills
- Amazon Web Services
- Amazon S3
- Big Data
- Google Cloud Platform
- Extract, Transform, Load
- Python
- PySpark
- Snow Flake Schema
Summary
Job Title: Senior Data Engineer – Cloud, Big Data & Data Platforms
Location: Seattle, WA (Onsite)
Duration: Long term
Experience: 5+ Years
Job Summary
We are looking for a Senior Data Engineer with 4+ years of experience in designing, developing, optimizing, and maintaining scalable data platforms and pipelines.
The ideal candidate should have strong hands-on experience with Python, PySpark, Apache Spark, SQL, Airflow, AWS/Google Cloud Platform, Data Lakehouse, ETL/ELT, and Data Warehousing.
The candidate should be comfortable working with both batch and streaming data processing, cloud-based data platforms, large-scale transformations, data migration, data quality, and production pipeline optimization. Experience with Databricks, Delta Lake, Kafka/Flink, Snowflake, BigQuery, Trino, Iceberg, or similar technologies will be an added advantage.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python, PySpark, Spark SQL, and SQL.
- Build high-performance batch and streaming data processing solutions using Apache Spark and technologies such as Kafka/Flink.
- Develop cloud-based data solutions across AWS and/or Google Cloud Platform.
- Work with AWS services including S3, EMR, Glue, Athena, Redshift, Lambda, and CloudWatch.
- Work with Google Cloud Platform services including Dataproc, GCS, BigQuery, and Cloud SQL.
- Develop and manage workflows using Apache Airflow, including DAG development, scheduling, monitoring, retries, and failure handling.
- Design and implement Medallion/Lakehouse architectures using Bronze, Silver, and Gold data layers.
- Work with Databricks, Delta Lake, Iceberg, Parquet, Trino, or equivalent data lakehouse technologies.
- Develop incremental processing, CDC, MERGE/UPSERT, schema evolution, and SCD Type I/II implementations.
- Design dimensional models involving fact and dimension tables, Star Schema, and Snowflake Schema.
- Build reliable data ingestion frameworks from databases, APIs, files, SaaS applications, and streaming sources.
- Optimize Spark workloads using partitioning, partition pruning, broadcast joins, AQE, caching, Z-Ordering, and shuffle optimization.
- Troubleshoot and optimize production pipelines for performance, scalability, SLA adherence, and cloud cost efficiency.
- Implement automated data quality, validation, reconciliation, monitoring, and alerting frameworks.
- Participate in cloud/platform migrations and perform data validation and production-readiness testing.
- Develop CI/CD and deployment workflows using Git, Jenkins, GitHub Actions, Docker, and Kubernetes.
- Collaborate with analytics, backend, product, and business teams to deliver reliable data products.
- Ensure data security, governance, lineage, compliance, and operational reliability across data platforms.
- Dice Id: 10110952
- Position Id: 9109023
- Posted 5 hours ago
Company Info
About GTSS Inc
"If you want to build something big, you have to start with a small step!"
Beginning 2003, GTSS has been a leader in providing flex staffing solutions for Fortune 500, SMBs and startups from a single resource to a full-fledged project team for their long term innovation programs. Given the scarcity of emerging technology skills and full stack developers we offer ‘uberisation of talent’ providing an effective and efficient staffing model for business to be agile and build teams as and when they want.


Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs