Job Title: Google Cloud Platform Data Engineer
Location: San Francisco, CA Hybrid / Local Candidates only
Employment Type: Full-Time / Contract
Experience : 7+
Job Summary
We are seeking a skilled Google Cloud Platform Data Engineer with strong hands-on experience designing, developing, and maintaining scalable data pipelines and cloud-based data platforms on Google Cloud Platform (Google Cloud Platform).
The ideal candidate will have proven experience with Python, SQL, PySpark, BigQuery, Apache Airflow, and dbt, along with a strong understanding of modern ETL/ELT frameworks, data warehousing, distributed data processing, and cloud data engineering practices.
The successful candidate will work closely with Data Analysts, Data Scientists, Business Stakeholders, and Engineering teams to build reliable, scalable, and high-performing data solutions.
Key Responsibilities
Design, develop, and maintain scalable data pipelines using Python and PySpark.
Build, maintain, and optimize ETL/ELT workflows using Apache Airflow.
Develop data transformation models and workflows using dbt.
Design, develop, and optimize data warehouse solutions using Google BigQuery.
Write efficient and optimized SQL queries for large-scale data processing and analytics.
Work with Google Cloud Platform services to build reliable and scalable cloud data solutions.
Ensure data quality, integrity, accuracy, and reliability across data pipelines.
Collaborate with Data Analysts, Data Scientists, Business Stakeholders, and Engineering teams to understand data requirements.
Monitor, troubleshoot, and resolve production data pipeline and workflow issues.
Implement best practices for logging, monitoring, testing, and CI/CD.
Optimize Google Cloud Platform cloud resource utilization and data processing costs.
Participate in code reviews and follow engineering best practices.
Contribute to the design and implementation of scalable data architecture and data engineering solutions.
Required Skills & Experience
Strong hands-on experience with Python and SQL.
Hands-on experience with PySpark and distributed data processing.
Strong experience with Google BigQuery.
Hands-on experience with Apache Airflow for data pipeline orchestration.
Hands-on experience with dbt for data transformation and data modeling.
Strong understanding of ETL/ELT frameworks.
Strong understanding of data warehousing concepts.
Experience working with Google Cloud Platform (Google Cloud Platform) and its data services.
Experience with Google Cloud Platform services including:
BigQuery
Cloud Storage (GCS)
Dataproc
Cloud Composer
Pub/Sub
Experience with Git and CI/CD practices.
Strong analytical, troubleshooting, and problem-solving skills.
Good to Have
Experience building streaming data pipelines using Apache Kafka or Google Pub/Sub.
Understanding of data governance, data security, and access control concepts.
Experience with Terraform or Infrastructure as Code (IaC).
Knowledge of Docker and Kubernetes.
Experience with cloud cost optimization and resource management.
Preferred Qualifications
Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
Google Cloud Platform certification is a plus.
Experience working in enterprise-scale data environments is preferred.