We are hiring for Epic Data Platform Lead at Santa Clara, CA onsite
JOB DESCRIPTION
EPIC Data Platform Lead
Databricks | Data Engineering and Analysis | Cloud Security | Governance | Encryption | Multi-Tenant Platforms
Function | Data Platforms / Advanced Analytics |
Role type | Technical Lead / Solution Lead |
Primary platform | Databricks Lakehouse on cloud |
Scope | EPIC data platform, dedicated tenant and multi-tenant capabilities |
Location | Santa Clara, CA |
Reporting relationship | To be determined |
Role Purpose
The EPIC Data Platform Lead will own the technical direction and implementation leadership for a secure, governed, scalable cloud data platform built on Databricks. The role combines hands-on data engineering and analytical problem solving with architecture leadership across tenant isolation, data governance, identity and access, encryption, observability, production readiness, and platform operations. The lead will translate business and engineering requirements into implementable platform capabilities for internal, customer-dedicated, and controlled multi-tenant use cases.
Expected Outcomes
01 | 02 | 03 |
Trusted data products Curated, traceable, analytics-ready data with clear ownership and quality controls. | Secure tenant boundaries Validated isolation across workspace, catalog, storage, identity, network, jobs, APIs, and exports. | Production-grade operations Observable, supportable, cost-aware services with automated deployment and evidence-based controls. |
Key Responsibilities
Platform architecture and technical leadership: Define target architecture, engineering standards, roadmaps, decision records, reusable patterns, and non-functional requirements for EPIC Databricks environments. Lead design reviews and make trade-offs across performance, security, operability, scalability, and cost.
Databricks implementation: Lead workspace, Unity Catalog, Delta Lake, pipeline, workflow, SQL warehouse, compute policy, external location, storage credential, and deployment-pattern implementation. Establish maintainable medallion-layer processing and production engineering practices.
Data engineering and analysis: Design and review batch, streaming, and event-driven ingestion; transformation and source-to-target logic; reconciliation; data profiling; exploratory analysis; root-cause analysis; and analytical data products. Use data to validate latency, completeness, linking, accuracy, and business-rule outcomes.
Security by design: Partner with cybersecurity, IAM, cloud, network, and application teams to implement least privilege, SSO/federation, service principals, secrets management, private connectivity, controlled egress, hardening, vulnerability remediation, and auditable access.
Governance and data protection: Implement data classification, taxonomy, ownership, metadata, lineage, retention, access reviews, fine-grained permissions, row filters, column masks, controlled sharing, DLP-aligned controls, and evidence-driven compliance.
Encryption and key management: Design and implement encryption in transit and at rest, customer-managed keys and BYOK patterns where required, KMS/HSM integration, key scope and separation, rotation, revocation, monitoring, recovery, and control validation.
Dedicated and multi-tenant delivery: Define tenant onboarding, registry, provisioning, configuration, isolation, routing, metering, offboarding, and migration patterns. Prevent unauthorized cross-tenant access and validate isolation through automated negative testing and periodic control reviews.
Observability and operations: Implement end-to-end logging, auditability, lineage, data-quality monitoring, health dashboards, alerting, SIEM integration, incident response, runbooks, service-level measures, capacity planning, and cost showback.
Delivery leadership: Own backlog quality, milestones, dependencies, risk mitigation, release readiness, production cutover, operational handoff, and stakeholder communication. Mentor engineers and coordinate delivery across data, cloud, security, governance, QA, infrastructure, and application teams.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, Data Science, or a related field, or equivalent practical experience.
- Strong experience leading the design and implementation of enterprise cloud data platforms, with substantial hands-on Databricks experience.
- Strong working knowledge of Apache Spark, Delta Lake, Databricks Workflows, Unity Catalog, SQL, and Python. Scala experience is beneficial.
- Demonstrated ability to perform complex data analysis, profiling, reconciliation, debugging, performance analysis, and root-cause investigation using large datasets.
- Experience implementing production-grade batch, streaming, micro-batch, event, and file-based ingestion patterns, including schema evolution, replay, backfill, and idempotent processing.
- Experience with cloud-native data services, object storage, IAM, private networking, key management, logging, monitoring, infrastructure as code, and CI/CD. AWS experience is preferred; Azure or Google Cloud Platform experience is also relevant.
- Strong knowledge of data security and governance concepts, including least privilege, identity federation, service principals, classification, lineage, retention, masking, audit logging, DLP, controlled data sharing, and regulated or customer-sensitive data handling.
- Experience designing or operating dedicated-tenant, multi-tenant, or customer-isolated platforms, including tenant lifecycle, logical and physical isolation, resource governance, and cross-tenant security testing.
- Experience implementing encryption in transit and at rest, cloud KMS or HSM integrations, customer-managed keys, BYOK, key rotation, separation of duties, and cryptographic control evidence.
- Strong architecture, technical documentation, stakeholder management, and engineering leadership skills, with the ability to convert ambiguous requirements into executable designs and delivery plans.
Preferred Qualifications
- Experience with Databricks on AWS, including S3, KMS, PrivateLink, VPC endpoints, IAM roles, CloudTrail, CloudWatch, and enterprise network controls.
- Experience with Kafka or equivalent event-streaming platforms and cloud edge or integration services.
- Experience with Databricks Asset Bundles, Terraform, Git-based workflows, automated testing, release pipelines, policy as code, and environment promotion.
- Experience building data-quality frameworks, lineage, observability, operational dashboards, and cost or usage reporting by tenant.
- Experience with Delta Sharing, APIs, BI tools, MLflow, AI/ML workloads, or governed data-product consumption patterns.
- Knowledge of Zero Trust, NIST CSF, CIS Controls, security architecture reviews, threat modeling, penetration testing, exception management, and audit evidence practices.
- Experience in semiconductor manufacturing, R&D, lab, metrology, equipment telemetry, or OT-integrated data environments is an advantage.
- Relevant Databricks, cloud architecture, data engineering, security, or governance certifications are desirable.
Core Technical Competencies
Competency | Expected depth | Evidence of capability |
Databricks and lakehouse | Expert | Spark, Delta Lake, Unity Catalog, Workflows, SQL, access patterns, performance, operations |
Data engineering and analysis | Expert | Ingestion, transformation, profiling, reconciliation, data quality, root-cause analysis, SQL/Python |
Security and governance | Advanced | IAM, least privilege, classification, lineage, masking, DLP, audit, SIEM, controlled sharing |
Encryption and key management | Advanced | TLS, encryption at rest, KMS/HSM, CMK/BYOK, rotation, revocation, evidence |
Tenant architecture | Advanced | Dedicated and multi-tenant patterns, isolation, provisioning, lifecycle, metering, testing |
Cloud and DevSecOps | Advanced | Private networking, infrastructure as code, CI/CD, secrets, observability, reliability, cost |
Leadership and delivery | Advanced | Architecture governance, planning, risk, production readiness, mentoring, stakeholder alignment |
Leadership Behaviors
- Customer and data protection mindset: treats tenant isolation, confidentiality, and evidence as design requirements, not afterthoughts.
- Outcome ownership: drives work from architecture and backlog through production validation and operational handoff.
- Analytical rigor: challenges assumptions, quantifies gaps, validates results, and communicates findings with traceable evidence.
- Pragmatic architecture: balances enterprise standards with delivery speed, maintainability, scalability, and supportability.
- Cross-functional influence: aligns engineering, security, governance, infrastructure, product, and business stakeholders without losing technical depth.
- Engineering excellence: promotes reusable components, automation, documentation, testing, peer review, and continuous improvement.
Measures of Success
- Architectural decisions, security controls, and governance requirements are translated into an executable and traceable engineering backlog.
- Databricks pipelines and analytical data products meet agreed quality, latency, reliability, security, and support readiness criteria.
- Dedicated and multi-tenant environments demonstrate validated isolation across data, identity, compute, storage, network, and operational boundaries.
- Encryption, key-management, access-governance, and audit controls are implemented with reviewable evidence and defined operational ownership.
- Platform releases include automated testing, observability, rollback, runbooks, production support, and stakeholder-ready status reporting.
- The engineering team becomes more effective through technical guidance, reusable patterns, stronger documentation, and disciplined delivery practices.
Candidate Evaluation Focus
Recruiting note: This role should be evaluated as a hands-on technical leadership position. Candidates should demonstrate both platform implementation depth and the ability to lead cross-functional delivery, rather than architecture-only or people-management-only experience.
Assessment area | Recommended evidence |
Databricks depth | Architecture walkthrough plus hands-on discussion of Unity Catalog, Delta design, performance, pipelines, compute policies, deployment, and production operations. |
Data analysis | Case exercise requiring SQL/Python reasoning, reconciliation, anomaly investigation, data-quality diagnosis, and clear communication of findings. |
Security and governance | Scenario covering IAM, private connectivity, egress, classification, lineage, masking, access review, audit logging, and exception handling. |
Encryption | Design discussion covering CMK/BYOK, KMS/HSM, key hierarchy, tenant key separation, rotation, revocation, recovery, and evidence. |
Multi-tenancy | Threat and architecture review for tenant onboarding, isolation boundaries, metadata-driven routing, cross-tenant negative testing, observability, and offboarding. |
Leadership | Examples of driving ambiguous platform work, resolving cross-team dependencies, making trade-offs, mentoring engineers, and achieving production readiness. |