Data Engineer
Location: Indianapolis, IN (onsite 5 days per week)
Exp: 9+
Project Details:
Pharmaceutical’s client is building a governed data and AI platform, integrating device, laboratory, partner, and document data into a unified foundation that supports regulatory reporting, advanced analytics, and AI-driven scientific insights
The Data Engineer is a hands-on builder responsible for developing data pipelines, API integrations, and AI infrastructure that bring structured and unstructured data into a governed Azure-based architecture. This delivery-focused role requires designing, coding, testing, and maintaining production-ready solutions across both AI document ingestion and structured ETL/ELT data engineering tracks. This role integrates structured and unstructured data from laboratory systems, product lifecycle applications, and manufacturing partners into a governed Azure-based data platform, creating a scalable, audit-ready digital thread that supports analytics, AI, and regulatory compliance.
Key Responsibilities
· Design, build, and maintain production ETL/ELT pipelines integrating laboratory and operational systems (e.g., Darwin, Teamcenter/PLM, LabVantage LIMS, Jama, TurboAC, and Qdocs/Veeva) into Microsoft Azure Fabric Lakehouse and PostgreSQL.
· Develop and optimize Bronze, Silver, and Gold medallion architecture, including schema mapping, data modeling, referential integrity, and performance optimization.
· Build scalable API integrations and cross-cloud data pipelines across Azure and AWS to support enterprise data integration.
· Implement automated data quality controls, controlled vocabulary normalization, schema validation, Q-gate/specification checks, data lineage, and audit trails to ensure GxP and ALCOA+ compliance.
· Monitor, troubleshoot, and optimize pipeline performance, reliability, error handling, and operational monitoring in production environments.
· Collaborate with business, engineering, and IT teams to integrate data sources and establish a governed, scalable digital thread supporting analytics, AI, and regulatory reporting.
MUST HAVE Experience:
· 5+ years of hands-on Data Engineering experience building production ETL/ELT pipelines and API integrations.
· Expert-level Python and/or PySpark for data ingestion, transformation, orchestration, testing, and CI/CD.
· Strong experience with Microsoft Azure Fabric (Lakehouse, Data Factory, Fabric Pipelines, Delta Lake) delivering end-to-end production solutions.
· Hands-on AWS experience, including S3, Glue (or equivalent), and RDS/Aurora.
· Strong SQL and PostgreSQL experience, including normalized schema design, query optimization, indexing, and performance tuning.
· Experience with data modeling, medallion architecture, and Lakehouse design patterns.
· Experience implementing data lineage, quality controls, schema validation, error handling, and monitoring in regulated environments.
· Knowledge of GxP, GMP, Google Cloud Platform within pharmaceutical, biotechnology, or medical device environments.(21 CFR Part 11)
· SCM, WMS, PLM, MM, QA, any validated applications with in a FDA Regulated systems