MUST-HAVE SKILLS & EXPERIENCE
Programming & Data Engineering
Expert - level proficiency in Python, Scala, and PySpark, with a strong track record of designing and delivering production-ready, modular, and well-tested solutions; developing and troubleshooting Spark workloads; and optimizing large-scale batch and streaming data pipelines using Delta Lake and Spark technologies.
Strong SQL and data modelling - dimensional and normalised; schema design and data contract definition.
Databricks expertise - Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving.
Azure data stack - ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks, error handling), Azure Event Hubs.
AI & Machine Learning
3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, structured output, chunking and embedding strategy, vector and hybrid retrieval, and prompt engineering.
Evaluation discipline - golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback loops; you measure AI quality, not assert it.
Hands-on experience with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
Metadata-driven thinking - schema inference, data profiling, lineage, catalogs, and configuration-driven frameworks that onboard the next source without new code.
Architecture & Governance
12 18 years of total experience in data engineering, data platform delivery, or related disciplines.
Proven delivery of a medallion / lakehouse architecture at enterprise scale - not just familiarity with the concept.
Azure security and governance - Entra ID, managed identities, RBAC, POSIX ACLs on ADLS Gen2, Key Vault, private endpoints, and PII handling.
CI/CD and infrastructure as code - Azure DevOps, Terraform, Databricks Asset Bundles, and automated testing of data pipelines.
Clear technical writing and the ability to present and defend a design to both engineers and non-technical stakeholders.
STRONGLY PREFERRED
Knowledge graphs and ontologies: RDF/SPARQL, property graphs (Neo4j), or graph modelling over a lakehouse.
Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale, including access control and ambiguity handling.
ML-based anomaly detection on time-series or transactional financial data.
Financial services or insurance domain knowledge: finance close, general ledger, subledger, reconciliation, or actuarial data.
LLMOps and MLOps: model versioning, prompt versioning, cost governance, and observability tooling.
Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification.
dbt, Great Expectations, or similar data-quality and transformation tooling.
Workday, Prism, or Accounting Center exposure.