Lead Data Engineer

Hybrid in Cary, NC, US • Posted 1 hour ago • Updated 1 hour ago
Full Time
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • ADF
  • Access Control
  • Actuarial Science
  • Apache Kafka
  • Apache Spark
  • Artificial Intelligence
  • Automated Testing
  • Batch File
  • Change Data Capture
  • Continuous Delivery
  • Continuous Integration
  • DML
  • Data Engineering
  • Data Modeling
  • Data Processing
  • Data Profiling
  • Data Quality
  • Database
  • Databricks
  • Debugging
  • DevOps
  • Documentation
  • Evaluation
  • Finance
  • Financial Services
  • GL
  • Information Security Governance
  • Insurance
  • LangChain
  • Layout
  • Leadership
  • LlamaIndex
  • Machine Learning (ML)
  • Machine Learning Operations (ML Ops)
  • Management
  • Mapping
  • Mentorship
  • Meta-data Management
  • Microsoft Azure
  • Natural Language
  • Neo4j
  • Ontologies
  • POSIX
  • Performance Tuning
  • Prism
  • Privacy
  • Prompt Engineering
  • Public Relations
  • PySpark
  • Python
  • RBAC
  • Regression Analysis
  • Resource Description Framework
  • SPARQL
  • SQL
  • Scala
  • Semantics
  • Shipping
  • Streaming
  • Technical Drafting
  • Technical Writing
  • Terraform
  • Time Series
  • Unity
  • Workday
  • Workflow

Summary

Role: Lead Data Engineer (Hands-On)
Location: Cary, NC (On-site / Hybrid)
Experience: 12-18 years
Employment: Full-Time
Salary: $140K - $145K per annum plus benefits
Eligibility: ONLY W2, NO C2C

ABOUT THE ENGAGEMENT
A centralized, AI-first enterprise Data Hub for a global insurance and
financial services client on Azure Databricks. The platform ingests
150+ inbound data feeds, distributes to 35+ downstream systems, and is
organized as a medallion architecture (Bronze / Silver / Gold). AI is
embedded in ingestion, canonical mapping, data quality,
reconciliation, and business user access from day one.

This is a senior hands-on leadership role. The candidate will own the
end-to-end technical design of the data and AI layers, build reference
implementations for the engineering team, and ship production-grade
Python, Scala, and PySpark code every week. Candidates who have not
written or reviewed production code in the past year are not a fit.

WHAT THE ROLE OWNS
- Data platform architecture and engineering: lakehouse architecture
(Bronze / Silver / Gold contracts, ADLS Gen2 zone layout, Delta Lake
table design, partitioning, schema evolution, retention).
- Metadata-driven, parameterized ingestion frameworks for batch files,
database extracts, CDC feeds and streaming (Azure Event Hubs / Kafka,
Spark Structured Streaming).
- Canonical PySpark and Scala Spark jobs, coding and testing
standards, PR reviews, production incident debugging, Spark cluster
tuning and cost guardrails.
- CI/CD for Databricks and ADF in Azure DevOps using Databricks Asset
Bundles and Terraform; observability with Azure Monitor and Log
Analytics.
- AI-augmented ingestion and canonical mapping: auto-generated bridge
documents, DML, canonical table definitions; AI-assisted
source-to-canonical mapping with human review gate.
- AI-driven data quality, anomaly detection (data drift, schema drift,
volume shifts, reconciliation breaks), automated reconciliation, and
synthetic privacy-preserving test data.
- Semantic layer and knowledge graph, plus a GPT-powered
conversational interface (text-to-SQL / semantic-layer retrieval) with
row- and column-level security.
- Governance and leadership: Unity Catalog (lineage, access control,
PII standards), Architecture Review Boards and AI governance forums,
mentoring engineers, documentation.

MUST-HAVE SKILLS & EXPERIENCE
- Expert-level Python, Scala and PySpark: production-ready, modular,
well-tested solutions; Spark workload troubleshooting; optimizing
large-scale batch and streaming pipelines using Delta Lake.
- Strong SQL and data modelling (dimensional and normalised), schema
design, data contracts.
- Databricks expertise: Delta Lake, Unity Catalog, Jobs & Workflows,
cluster and pool management, performance tuning, Model Serving.
- Azure data stack: ADLS Gen2 (zone design, ACLs, lifecycle), Azure
Data Factory (parameterized / metadata-driven frameworks), Azure Event
Hubs.
- 3+ years designing and shipping LLM-based systems in production: RAG
pipelines, agentic / tool-calling workflows, chunking and embedding
strategy, vector and hybrid retrieval, prompt engineering.
- Evaluation discipline: golden datasets, regression suites, accuracy
and hallucination tracking, human-in-the-loop feedback.
- Hands-on with LangChain, LlamaIndex or LangGraph, plus at least one
provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
- Metadata-driven frameworks: schema inference, data profiling,
lineage, catalogs.
- 12-18 years of total experience in data engineering / data platform delivery.
- Proven enterprise-scale delivery of a medallion / lakehouse architecture.
- Azure security and governance: Entra ID, managed identities, RBAC,
POSIX ACLs, Key Vault, private endpoints, PII handling.
- CI/CD and IaC: Azure DevOps, Terraform, Databricks Asset Bundles,
automated testing of data pipelines.
- Clear technical writing and ability to present and defend designs to
engineers and non-technical stakeholders.

STRONGLY PREFERRED
- Knowledge graphs and ontologies (RDF/SPARQL, Neo4j, graph modelling
over a lakehouse).
- Text-to-SQL or semantic-layer-backed natural-language query systems
at enterprise scale.
- ML-based anomaly detection on time-series or transactional financial data.
- Financial services or insurance domain (finance close, GL,
subledger, reconciliation, actuarial data).
- LLMOps / MLOps: model and prompt versioning, cost governance, observability.
- Databricks Data Engineer Professional, Azure DP-203 / DP-700, or
AZ-305 certification.
- dbt, Great Expectations or similar; Workday, Prism or Accounting
Center exposure.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10113363
  • Position Id: 9107981
  • Posted 1 hour ago
Contact the job poster
Abhishek Sharma

Abhishek Sharma

Senior Recruiter @ Innovative Information Technologies, Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Cary, North Carolina

•

Today

Easy Apply

Full-time, Part-time, Contract, Third Party

Compensation information provided in the description

Hybrid in Cary, North Carolina

•

Today

Easy Apply

Full-time

130000 - 140000

Cary, North Carolina

•

Yesterday

Easy Apply

Full-time, Third Party

$130000 - $140000

Hybrid in Cary, North Carolina

•

Today

Easy Apply

Full-time, Third Party

Depends on Experience

Search all similar jobs