Role Summary
As a Data Engineer you ll design and build the data infrastructure that powers analytics, AI, and client-facing platforms across our portfolio of engagements. You ll sit at the center of our Data, Analytics & AI practice working alongside data scientists, analytics engineers, and solution architects to turn raw, complex data into reliable, production-grade pipelines. This isn t a support role. The pipelines you build directly shape what our clients can see, measure, and act on.
The Impact You ll Have
Our clients are large, complex organizations manufacturers, distributors, financial services firms that sit on enormous volumes of data they can t yet fully use. You ll help change that. By building and maintaining Databricks-based lakehouse architectures, you ll enable the kind of clean, governed, query-ready data layers that unlock real analytics programs and AI use cases not proofs of concept, but production solutions.
The work here touches both depth and breadth. On any given engagement, you might be designing Delta Live Tables pipelines for a client s demand forecasting model, optimizing Unity Catalog configurations for cross-team data access, or collaborating with our analytics team to ensure a reporting layer performs at scale. The variety is real, and so is the ownership.
you ll contribute to a growing Data, Analytics & AI practice that s building repeatable delivery accelerators and platform capabilities. Your work feeds directly into how we expand what s possible for clients and how we sharpen our own technical edge as an organization.
What You ll Do Pipeline Design & Development
Build and maintain scalable data pipelines using Apache Spark and Databricks, including batch and streaming workloads
Develop Delta Live Tables (DLT) workflows that automate data quality checks and reduce time-to-insight for analytics teams
Translate complex business data requirements often surfaced by strategy or client services into well-structured, documented pipeline logic
Lakehouse Architecture
Design and implement Delta Lake architectures including Bronze/Silver/Gold medallion patterns for client environments
Configure and manage Unity Catalog for fine-grained access control, data lineage, and cross-workspace governance
Optimize storage formats, partition strategies, and compute configurations to control cost and improve query performance
Data Modeling & Transformation
Build and maintain data models that serve downstream analytics, dashboards, and ML feature pipelines
Collaborate with analytics engineers and data scientists to align transformation logic with reporting and model training requirements
Apply dbt or equivalent transformation tooling where it fits the client s stack and delivery pattern
Data Quality & ObservabilityImplement data validation, schema enforcement, and monitoring frameworks to catch issues before they reach production
Document data lineage, transformation logic, and pipeline dependencies to support governance and audit requirements
Participate in incident response and root cause analysis when pipeline failures or data quality issues arise
Client Delivery & Cross-Functional Collaboration
Work closely with solution architects, data scientists, and delivery leads to scope, estimate, and execute data engineering work within client engagements
Communicate technical design decisions and tradeoffs clearly to both internal team members and client stakeholders
Contribute to reusable frameworks, templates, and delivery accelerators that raise the quality bar across the practice
What You ll Need
3 5 years of experience in data engineering, with a track record of delivering pipelines in production environments
Hands-on Databricks experience: Delta Lake, Databricks Workflows, Unity Catalog, and Databricks SQL
Strong proficiency in Apache Spark (PySpark preferred) and SQL for large-scale data transformation
Experience with cloud data platforms AWS (S3, Glue, Redshift, Lambda) and/or Azure (ADLS, Synapse, ADF)
Familiarity with Delta Live Tables or equivalent declarative pipeline frameworks
Solid understanding of data modeling concepts dimensional modeling, medallion architecture, schema design
Experience working in Agile or sprint-based delivery environments
Comfortable working directly with stakeholders to translate requirements into technical solutions
Bachelor s degree in Computer Science, Engineering, Information Systems, or equivalent practical experience
Future-Ready Skills (Nice to Have)
Experience with Databricks Mosaic AI or MLflow for feature engineering and model tracking in support of ML pipelines
Familiarity with AI-enabled data quality or pipeline observability tooling (e.g., Monte Carlo, Anomalo, or similar)
Exposure to data governance platforms such as Collibra, Alation, or Microsoft Purview
Experience in a consulting, agency, or managed services environment where you ve navigated multiple client contexts
Working knowledge of streaming architectures using Apache Kafka or Databricks Structured Streaming