Job Title: Lead Data Engineer (Hands-On)
Job Location: Cary, NC (On-site / Hybrid)
Job Duration: Full-Time
Experience: 12-18 years
About The Engagement
This is a senior hands-on leadership role. You will own the end-to-end technical design of the data and AI layers, build the reference implementations your engineers work from, and ship production-grade Python, Scala, and PySpark code every week. If you have not written or reviewed production code in the past year, this is not the right fit.
What You Will Own
Data Platform Architecture & Engineering
- Own the end-to-end lakehouse architecture: Bronze / Silver / Gold layer contracts, zone layout on ADLS Gen2, Delta Lake table design, partitioning, schema evolution, and retention strategy.
- Design and build metadata-driven, parameterized ingestion frameworks that onboard new data sources without bespoke pipeline code for every feed.
- Write canonical PySpark and Scala Spark transformation jobs that serve as the team reference; set coding and testing standards, review pull requests, and debug production incidents.
- Design for scale and cost: tune Spark clusters and pools, apply partition pruning and caching strategies, and set cost guardrails as data volumes grow.
- Build and automate CI/CD for Databricks and ADF pipelines in Azure DevOps using Databricks Asset Bundles and Terraform; maintain platform observability with Azure Monitor and Log Analytics.
- Design repeatable patterns for batch files, database extracts, CDC feeds, and streaming ingestion using Azure Event Hubs / Kafka and Spark Structured Streaming.
AI-Augmented Ingestion & Canonical Mapping
- Design and build the AI-augmented metadata ingestion framework that auto-generates bridge documents, DML statements, canonical table definitions, and control metadata from source schemas.
- Build AI-assisted source-to-canonical attribute mapping: schema reasoning, data profiling, confidence scoring, and a human review / approval gate before any mapping reaches production.
- Generate file-level and record-level validation rules from historical data and metadata analysis, feeding a configurable, rules-engine-backed data quality framework.
AI-Driven Data Quality, Anomaly Detection & Testing
- Build AI-assisted data quality that analyses patterns across Bronze, Silver, and Gold layers to propose DQ rules beyond predefined checks.
- Deliver anomaly detection covering outliers, data drift, schema drift, volume shifts, and reconciliation breaks — with actionable alerting, not noise.
- Build AI-assisted automated reconciliation and test-data generation to feed the platform's automated testing framework.
- Produce synthetic, privacy-preserving datasets for lower environments using differential privacy, format-preserving masking, and referential-integrity-safe generation.
Semantic Layer & Conversational Data Access
- Design and build an ontology-driven semantic layer and knowledge graph modelling relationships across finance data entities, powering data discovery, semantic integration, and AI/BI tooling.
- Build a GPT-powered conversational interface for natural-language querying of financial data: text-to-SQL or semantic-layer-mediated retrieval grounded in the knowledge graph, with row-level and column-level security enforced and every answer traceable to source.
Governance, Architecture Reviews & Team Leadership
- Own end-to-end AI architecture decisions: model selection, RAG and retrieval design, prompt strategy, evaluation harnesses, guardrails, cost and latency budgets, and observability.
- Implement Unity Catalog for cataloguing, lineage, and fine-grained access control; define PII classification, masking, tokenization, and encryption standards across every layer.
- Take designs through Architecture Review Boards and AI governance forums, covering responsible AI, data residency, model approval, auditability, and human-in-the-loop controls.
- Mentor data engineers, run design reviews, and set the engineering patterns the team builds on — without becoming a bottleneck.
- Produce documentation and reusable components good enough for the client's team to operate the platform independently at engagement end.
Must-Have Skills & Experience
Programming & Data Engineering
- Expert - level proficiency in Python, Scala, and PySpark, with a strong track record of designing and delivering production-ready, modular, and well-tested solutions; developing and troubleshooting Spark workloads; and optimizing large-scale batch and streaming data pipelines using Delta Lake and Spark technologies.
- Strong SQL and data modelling — dimensional and normalised; schema design and data contract definition.
- Databricks expertise — Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving.
- Azure data stack — ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks, error handling), Azure Event Hubs.
AI & Machine Learning
- 3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, structured output, chunking and embedding strategy, vector and hybrid retrieval, and prompt engineering.
- Evaluation discipline — golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback loops; you measure AI quality, not assert it.
- Hands-on experience with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
- Metadata-driven thinking — schema inference, data profiling, lineage, catalogs, and configuration-driven frameworks that onboard the next source without new code.
Architecture & Governance
- 12–18 years of total experience in data engineering, data platform delivery, or related disciplines.
- Proven delivery of a medallion / lakehouse architecture at enterprise scale — not just familiarity with the concept.
- Azure security and governance — Entra ID, managed identities, RBAC, POSIX ACLs on ADLS Gen2, Key Vault, private endpoints, and PII handling.
- CI/CD and infrastructure as code — Azure DevOps, Terraform, Databricks Asset Bundles, and automated testing of data pipelines.
- Clear technical writing and the ability to present and defend a design to both engineers and non-technical stakeholders.
Strongly Preferred
- Knowledge graphs and ontologies: RDF/SPARQL, property graphs (Neo4j), or graph modelling over a lakehouse.
- Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale, including access control and ambiguity handling.
- ML-based anomaly detection on time-series or transactional financial data.
- Financial services or insurance domain knowledge: finance close, general ledger, subledger, reconciliation, or actuarial data.
- LLMOps and MLOps: model versioning, prompt versioning, cost governance, and observability tooling.
- Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification.
- dbt, Great Expectations, or similar data-quality and transformation tooling.
- Workday, Prism, or Accounting Center exposure.
ABOUT US
Apptad offers strategic consulting, enterprise information management and digital transformation services. With globally connected offices in US and India along with a team of trained and certified IT resources, Apptad ensures quick and effective delivery to its customers.Apptad is relentlessly reinventing the outlook of how companies leverage data.
With an effort to enable our customers the ability to solve biggest problems within their organization.We perceive our clients’ problems and respond with custom solutions instead of handing over boilerplate responses.
OUR MISSION
Customer Focus: We listen carefully to the needs of our clients so that we know what’s important for their business and can design a customized solution for their business.
Innovation: As a firm, we believe in constantly upgrading ourselves and improving our solutions to adapt to the changing landscape of technology.
Accountability and Ethics: We believe in taking our commitments as seriously as our customers and living up to them while building trust for a long term business relationship.