Senior Big Query & Spark Data Engineer

Remote • Posted 2 hours ago • Updated 2 hours ago
Contract W2
12 Months
No Travel Required
Remote
$54/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • GCP Bigquery with Spark

Summary

Job Summary

We are looking for an experienced Senior Data Engineer with strong hands-on expertise in Google BigQuery, Apache Spark, PySpark, Python, and SQL. The candidate will design and develop large-scale data pipelines, optimize distributed data-processing workloads, and build scalable analytical data platforms on Google Cloud.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT data pipelines using Python, SQL, Spark, and PySpark.
  • Build and optimize enterprise-scale data solutions using Google BigQuery as the analytical data warehouse.
  • Develop complex BigQuery SQL involving CTEs, window functions, nested/repeated fields, ARRAY/STRUCT operations, MERGE statements, and incremental processing.
  • Design efficient BigQuery tables using partitioning, clustering, materialized views, and appropriate data modeling techniques.
  • Analyze BigQuery query execution plans and optimize queries to reduce slot consumption, bytes scanned, execution time, and overall processing cost.
  • Implement incremental ingestion and transformation patterns rather than performing unnecessary full-table processing.
  • Develop large-scale distributed processing applications using Apache Spark and PySpark.
  • Work extensively with Spark DataFrames, Spark SQL, transformations, actions, joins, aggregations, and window operations.
  • Troubleshoot and optimize Spark workloads using Spark UI, execution plans, DAGs, stages, tasks, and executor metrics.
  • Perform advanced Spark performance tuning including partition management, repartition/coalesce strategies, predicate pushdown, partition pruning, caching/persistence, broadcast joins, and Adaptive Query Execution (AQE).
  • Identify and resolve data skew, shuffle bottlenecks, executor memory issues, excessive spills, and long-running stages.
  • Tune Spark configurations including executor memory, cores, shuffle partitions, serialization, and dynamic resource allocation based on workload requirements.
  • Design scalable processing patterns for multi-terabyte datasets while minimizing unnecessary data movement and shuffle operations.
  • Build batch and, where required, near-real-time data processing pipelines using appropriate Google Cloud Platform services.
  • Integrate BigQuery and Spark with services such as Google Cloud Storage, Dataproc, Dataflow, Pub/Sub, and Cloud Composer/Airflow.
  • Design dimensional and analytical data models including fact tables, dimension tables, star schemas, curated datasets, and reporting layers.
  • Implement CDC, incremental loads, deduplication, late-arriving data handling, SCD Type 1/Type 2, and idempotent pipeline patterns.
  • Build reusable frameworks for ingestion, transformation, validation, logging, exception handling, and pipeline monitoring.
  • Implement automated data-quality checks for completeness, uniqueness, accuracy, referential integrity, schema validation, and business-rule validation.
  • Troubleshoot production pipeline failures and perform root-cause analysis across Spark jobs, SQL workloads, source systems, and downstream datasets.
  • Implement monitoring and alerting for pipeline failures, SLA violations, data-quality issues, and abnormal processing behavior.
  • Work with structured, semi-structured, and large-volume datasets including JSON, Parquet, Avro, and CSV.
  • Apply security and governance practices including IAM, service accounts, BigQuery authorized views, row-level security, column-level security, and least-privilege access.
  • Participate in code reviews, technical design discussions, performance optimization, deployment, and production support.
  • Collaborate with analytics, BI, data science, and application teams to deliver reliable datasets for reporting, analytics, and AI/ML use cases.

Preferred Qualifications

  • 7–8+ years of overall Data Engineering experience.
  • Strong production experience with BigQuery and Apache Spark/PySpark.
  • Experience designing enterprise-scale cloud data platforms on Google Cloud Platform.
  • Strong understanding of distributed computing, Spark internals, and query optimization.
  • Experience processing datasets ranging from hundreds of gigabytes to multiple terabytes.
  • Strong understanding of data warehouse architecture and dimensional modeling.
  • Experience with CI/CD, Git, automated testing, and Infrastructure as Code is preferred.
  • Experience supporting analytics, Looker/BI, machine learning, or AI-oriented datasets is a plus.

Core Technology Stack

BigQuery | Apache Spark | PySpark | Python | SQL | Google Cloud Platform | Dataproc | GCS | Pub/Sub | Cloud Composer | Airflow | Parquet | Avro | Git | CI/CD | Data Modeling | ETL/ELT

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10113656
  • Position Id: 9081191
  • Posted 2 hours ago
Contact the job poster
AV

Aditya Varma

Recruiter @ Peritus Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

13d ago

Easy Apply

Full-time

120,000 - 200,000

Remote

7d ago

Easy Apply

Contract, Third Party

Depends on Experience

Remote

Today

Easy Apply

Contract

50 - 60

Remote or Minneapolis, Minnesota

Today

Easy Apply

Contract

Depends on Experience

Search all similar jobs