Position: Data Engineer
Location: San Jose, CA (Onsite)
Lead the end-to-end migration (transformation and load) of high-volume, high-complexity data into a new PostgreSQL-based infrastructure
Design, build, and maintain robust ETL/ELT pipelines capable of processing millions of records reliably and efficiently
Map and reconcile complex legacy data structures across multiple business entities into a unified target schema
Define and implement data validation frameworks to ensure integrity, completeness, and accuracy throughout the migration
Optimize pipeline performance, including query tuning, indexing strategy, and batch/incremental load design in PostgreSQL
Identify, troubleshoot, and resolve data quality issues, schema mismatches, and pipeline failures
Document data mappings, transformation logic, and migration runbooks for engineering and business stakeholders
Partner with business and technical stakeholders to align migration scope, timelines, and data requirements
Establish monitoring, logging, and alerting to track pipeline health and data quality post-migration
Mentor junior data engineers and contribute to engineering best practices and standards
Required Qualifications
5+ years of experience in data engineering, with demonstrated ownership of large-scale data migration projects
Strong to expert-level proficiency in Python for building and automating data pipelines
Deep hands-on experience with PostgreSQL/Oracle , including schema design, query optimization, and performance tuning
Proven experience designing and managing ETL/ELT pipelines at scale (millions of records)
Experience mapping and transforming complex, legacy data structures across disparate systems or business entities
Strong understanding of data validation, reconciliation, and quality assurance techniques
Solid grasp of data modeling principles (normalization, indexing, partitioning)
Experience with version control (Git) and CI/CD practices for data pipelines
Preferred Qualifications
Experience with orchestration tools (e.g., Airflow, Dagster, Prefect)
Familiarity with cloud data platforms (AWS, Google Cloud Platform, or Azure)
Experience with other relational or NoSQL databases and cross-database migrations
Background working in regulated or high-stakes data environments (finance, healthcare, etc.)
Experience with containerization (Docker) and infrastructure-as-code
Exposure to data quality/testing frameworks (e.g., Great Expectations, dbt tests)
What Success Looks Like
A fully migrated, validated dataset in PostgreSQL with zero critical data loss or corruption
ETL/ELT pipelines that are documented, repeatable, and optimized for ongoing operation
A clear audit trail of data lineage and validation results across all business entities involved