Overview
We are seeking an experienced Data Engineer with healthcare payer domain knowledge to design, build, and operate data pipelines and infrastructure within a large-scale, cloud-native data platform on Google Cloud Platform. The role operates in a complex, heterogeneous data environment spanning on-premises relational systems — including DB2 and Oracle — virtual integration layers, flat-file and SFTP-based sources, and real-time streaming feeds, all flowing into a Google Cloud Platform-based Lakehouse architecture.
The ideal candidate brings strong PySpark and Google Cloud Platform engineering skills, a deep understanding of data pipeline design for regulated healthcare data, and the ability to work effectively in environments where both legacy and modern technologies coexist. You will build and maintain the pipelines that ingest, transform, and deliver high-quality healthcare payer data to operational and analytical consumers at scale.
Primary Responsibilities
Pipeline Development
• Design, develop, and maintain scalable PySpark data pipelines on Dataproc for batch data processing across a wide range of healthcare payer source systems — including claims, member enrollment, provider, finance, and drug/pharmacy data originating from DB2, Oracle, Denodo, and flat-file SFTP sources.
• Build and maintain real-time streaming ingestion pipelines using Confluent Kafka, processing high-volume event streams for operational and analytical consumers.
• Implement Apache Iceberg table format across data lake storage layers, including schema management, partitioning design, compaction, and time-travel query support.
• Develop data transformation logic within curated and certified data layers — including cross-reference table processing, SCD1 pattern implementation, and history table management for member and provider data.
• Build and maintain orchestration workflows using Cloud Composer (Airflow) for batch pipeline scheduling, dependency management, SLA monitoring, and operational alerting.
Data Quality & Observability
• Implement automated data quality validation within pipelines — including record count checks, referential integrity validation, business rule enforcement, and anomaly detection — ensuring only certified data reaches serving and downstream consumers.
• Build monitoring and alerting frameworks for pipeline health, processing SLAs, and data freshness across batch and streaming workloads.
• Contribute to metadata and data lineage instrumentation within pipelines, supporting enterprise catalog and governance requirements.
• Participate in evaluations of data observability tools and quality frameworks as part of the evolving platform toolchain.
Platform Operations & Standards
• Implement and maintain security controls within data pipelines — including PHI/PII masking, column-level encryption, role-based access controls, and audit logging — in compliance with HIPAA, FedRAMP, and NIST standards.
• Contribute to Infrastructure-as-Code (IaC) practices for Dataproc cluster provisioning, Cloud Composer environment management, and Artifactory-based CI/CD deployment pipelines.
• Troubleshoot pipeline failures, optimize performance for high-volume workloads, and implement solutions for data reliability, idempotency, and retry handling.
• Produce technical documentation and contribute to knowledge transfer, ensuring pipeline logic and design decisions are well understood across the team.
Minimum Qualifications
• Bachelor''''s degree in Computer Science, Engineering, Data Science, or a related field.
• 7+ years of data engineering experience with at least 4 years building production-grade data pipelines for healthcare payer organizations covering claims, enrollment, provider, or analytics domains.
• Strong PySpark and SQL proficiency — the ability to write, optimize, and debug complex transformation logic in PySpark is a core requirement of this role.
• Hands-on experience with Google Cloud Platform data services: Dataproc, BigQuery, Cloud Storage (GCS), Pub/Sub, and Cloud Composer (Airflow).
• Experience working with on-premises relational source systems — including DB2 and Oracle — across ingestion, schema mapping, and incremental load patterns.
• Working knowledge of Apache Iceberg or Delta Lake, including understanding of open table format concepts, ACID transactions, schema evolution, and partition management.
• Experience with Apache Kafka or Confluent Kafka for real-time streaming data ingestion and event-driven pipeline design.
• Proficiency in Git-based version control and CI/CD practices for data pipeline deployment.
Preferred Experience
• Familiarity with virtual data integration platforms such as Denodo — understanding of how to consume data from virtual layers and how those patterns compare to direct source system ingestion.
• Experience with file-based ingestion from SFTP and cloud storage sources, including handling of Excel, CSV, text, and SAS file formats common in healthcare payer environments.
• Exposure to federated query and analytics platforms — such as Starburst — that operate across heterogeneous data sources including operational databases, cloud data warehouses, and Iceberg-based data lakes.
• Familiarity with modern data catalog and observability tools — such as Atlan or Monte Carlo — particularly in proof-of-concept or early adoption contexts.
• Understanding of Biglake Catalog or equivalent open catalog standards for unified metadata management across Google Cloud Platform-native and Iceberg-based storage.
• Knowledge of healthcare payer data structures and standards including X12 EDI, HL7, FHIR, and CMS reporting requirements.
• Strong problem-solving skills with the ability to diagnose and resolve complex data issues across heterogeneous, multi-layer pipeline architectures.