Pharma Data Engineer - Databricks AWS

Hybrid in Indianapolis, IN, US • Posted 4 hours ago • Updated 4 hours ago
Contract W2
12 Months
Occasional Travel Required
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Amazon Web Services
  • Data Management
  • Databricks
  • ELT
  • Amazon S3
  • GxP
  • SQL
  • Regulatory Compliance
  • Python
  • Pharmaceutics
  • AWS Glue
  • Datalake
  • AWS Athena
  • Orchestration
  • Airflow
  • Scala

Summary

Pharma Data Engineer – Databricks & AWS

Hybrid – Indianapolis, IN

We are seeking a Data Engineer with 3–5 years of experience working specifically within the pharma industry to join a pharma-focused data team. This is a senior-flavored engineering role that combines hands-on pipeline and platform work with significant business-facing responsibility — including translating business needs into technical specs, presenting to executive-level stakeholders, and helping stand up new data domains from the ground up. You will design, build, and govern the data infrastructure that powers analytics and reporting across the business, while also acting as a trusted technical partner to non-technical stakeholders.
Key Responsibilities

Design, build, and maintain scalable ETL/ELT pipelines (batch and streaming) using Databricks, AWS, and related orchestration tools

Write and optimize advanced SQL, and build data transformations in Python or Scala

Integrate external data sources via APIs and manage pipeline orchestration (Airflow, Databricks Workflows, AWS Glue)

Apply data quality, governance, cataloging, and lineage practices aligned with regulated-industry standards

Work within GxP-regulated data environments and apply awareness of data privacy/compliance considerations (e.g., 21 CFR Part 11, GDPR where applicable)

Partner with business stakeholders across the pharma value chain (R&D, Manufacturing & Quality, Commercial, Drug Development) to gather and translate requirements into technical specifications

Present technical work and data strategy to executive-level audiences

Prioritize high-impact data initiatives and proactively identify and avoid duplicated data efforts

Support change management and adoption of new data solutions across business teams. Help stand up new data domains from scratch (green-field build), not just maintain existing ones

Requirements

Required Qualifications
Data Engineering & Pipelines

  • Data Engineering & Pipelines
  • ETL/ELT development (batch and streaming)
  • Advanced SQL (joins, window functions, query optimization)
  • Python or Scala for data transformation
  • Data pipeline orchestration (Airflow, Databricks Workflows, AWS Glue)
  • API integration for external data source ingestion

Platforms & Tools

  • Databricks (Delta Lake, Unity Catalog, Genie)
  • Cloud platforms — AWS (S3, Glue, Athena) and/or Azure/Google Cloud Platform equivalents
  • Data warehousing concepts (dimensional modeling, star schema)
  • BI/visualization tools (Tableau, Power BI, or similar) to understand downstream consumption

Data Quality & Governance

  • Data profiling and cleansing techniques
  • Metadata management and data cataloging
  • Master data management (MDM) principles
  • Data lineage tracking
  • Data governance frameworks (especially regulated-industry standards)

Pharma / Life Sciences Domain Knowledge

  • Familiarity with GxP-regulated data environments
  • Understanding of the pharma value chain (R&D, Manufacturing & Quality, Commercial, Drug Development)
  • Awareness of data privacy/compliance considerations (21 CFR Part 11, GDPR where applicable)
  • Knowledge of common pharma data domains (clinical, manufacturing, quality, commercial)

Stakeholder Management

  • Requirements gathering and translation (business need → technical spec)
  • Cross-functional communication (Business ↔ IT)
  • Executive-level presentation skills (given EC visibility)
  • Change management / adoption support

Analytical & Strategic Thinking

  • Prioritization frameworks (identifying high-impact vs. low-value data asks)
  • Cost-avoidance mindset (spotting duplication before it happens)
  • Ability to work with ambiguity and evolving priorities

Project & Program Skills

  • Agile/Scrum familiarity
  • Documentation discipline (data dictionaries, source-to-target mappings)
  • Vendor/partner coordination (if external data sources are involved)

Nice-to-Have Differentiators

  • Prior consulting or client-facing delivery experience
  • Experience standing up new data domains from scratch (green-field vs. maintenance)
  • Familiarity with AI/GenAI-enabled analytics tools
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91133125
  • Position Id: 9073226
  • Posted 4 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Indianapolis, Indiana

Today

Easy Apply

Third Party, Contract

Depends on Experience

Remote

Today

Full-time

USD 149,000.00 - 248,000.00 per year

Remote or Texas

Today

Full-time

USD 110,000.00 - 200,000.00 per year

Remote

Today

Contract

$80.00 - $90.00

Search all similar jobs