Lead AI Engineer with Spark, AWS Services

Remote • Posted 1 hour ago • Updated 1 hour ago
Full Time
Remote
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Generative Artificial Intelligence (AI)
  • Quality Management
  • FOCUS
  • Clinical Trials
  • Vector Databases
  • Streaming
  • Extract
  • Transform
  • Load
  • ELT
  • Management
  • Data Quality
  • Artificial Intelligence
  • Machine Learning (ML)
  • Data Engineering
  • Workflow
  • Python
  • Apache Spark
  • SQL
  • Unstructured Data
  • PDF
  • Amazon Web Services
  • Amazon S3
  • Step-Functions
  • API
  • Amazon DynamoDB
  • Docker
  • English
  • Pharmaceutics
  • Life Sciences
  • Snow Flake Schema
  • Database
  • Amazon SageMaker
  • Continuous Integration
  • Continuous Delivery
  • Jenkins
  • Git
  • Bitbucket
  • Terraform
  • CDISC
  • SDTM

Summary

We're looking for a Lead AI Engineer skilled in Spark and AWS Services to become part of the RBQM Production Pod within the program. In this position, you'll construct and sustain data pipelines that drive AI/GenAI applications supporting Risk-Based Quality Management for clinical trials. The primary focus of this role includes RAG document ingestion, vector indexing, and developing data APIs for AI applications. Responsibilities Architect and construct RAG document ingestion pipelines (chunking, embedding, vector indexing) to support clinical trial quality data Establish and oversee vector databases (AWS OpenSearch) to support RAG-driven AI workflows Create batch and streaming ETL/ELT pipelines from the ground up for unstructured clinical data (PDF, DOCX, clinical reports) Construct and expose data APIs that AI applications can consume Enhance chunking strategies, embedding generation, and retrieval performance within RAG architectures Oversee data quality, lineage, and governance across AI/ML data pipelines Set up and sustain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB) Partner with Data Scientists and Backend Developers as part of a unified pod team Requirements Minimum 7 years of practical, large-scale data engineering experience Strong background in RAG document ingestion pipelines (chunking, embedding, vector indexing) Skilled in using AWS OpenSearch as a vector database for RAG workflows High-level command of Python, along with SQL and Spark SQL Experience transforming unstructured data (PDF, DOCX) for use in RAG/LLM applications Working knowledge of AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, DynamoDB Understanding of Docker-based containerization Ability to develop custom pipelines from the ground up, going beyond simple configuration of pre-built services Proficiency in English at a B2+ level Nice to have Experience within the pharmaceutical or life sciences sector Exposure to Snowflake and Pinecone (as an alternative vector database) Understanding of SageMaker processing jobs Proficiency with CI/CD tools (Jenkins, Git/Bitbucket) and infrastructure-as-code tools (CDK or Terraform) Familiarity with clinical data standards (CDISC, ADaM, SDTM)
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10330481
  • Position Id: 260b2b67079090e78bbdf19a58bffac3
  • Posted 1 hour ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

•

Today

Full-time

Remote

•

5d ago

Easy Apply

Full-time

Depends on Experience

Remote

•

Today

Full-time

USD 117,000.00 - 146,000.00 per year

Remote

•

Today

Full-time

USD 194,600.00 - 361,400.00 per year

Search all similar jobs