Location: Tustin, CA Salary: $160,000.00 USD Annually - $170,000.00 USD Annually Description:
Senior Platform Engineer
Irvine, CA
Fulltime
We are seeking a Senior Platform Engineer to build, administer, automate, secure, and operate enterprise Databricks and cloud data-platform environments. The role is responsible for Databricks workspace administration, Unity Catalog governance, identity and access management, compute and cluster policies, infrastructure automation, CI/CD enablement, monitoring, reliability, cost management, and production support.
The ideal candidate combines hands-on Databricks administration with strong AWS or Azure platform engineering, Terraform, Python, Kubernetes, CI/CD, security, networking, observability, and Site Reliability Engineering practices.
Key Responsibilities
Databricks Administration and Platform Engineering
Administer Databricks accounts, workspaces, metastores, catalogs, schemas, external locations, storage credentials, connections, shares, recipients, and platform configurations.
Provision and manage development, test, staging, and production workspaces using standardized, repeatable patterns.
Configure workspace settings, repositories, jobs, notebooks, SQL warehouses, instance pools, job clusters, all-purpose compute, and serverless capabilities.
Define and enforce cluster policies, approved runtime versions, libraries, init scripts, autoscaling, tagging, and compute-usage standards.
Manage Databricks Runtime and platform upgrades, compatibility testing, release planning, maintenance windows, and rollback procedures.
Support Databricks Workflows, Delta Live Tables or Lakeflow Declarative Pipelines, Databricks SQL, MLflow, model registry, feature engineering, and structured-streaming services.
Troubleshoot workspace, permissions, connectivity, compute, storage, job execution, library, runtime, and performance issues.
Maintain administration standards, platform documentation, knowledge articles, support procedures, and operational runbooks.
Unity Catalog, Identity, Security, and Governance
Design and manage Unity Catalog metastores, catalogs, schemas, managed and external tables, volumes, external locations, storage credentials, and grants.
Implement user and group provisioning through single sign-on, SCIM, identity-provider integration, and enterprise directory services.
Administer account-level and workspace-level users, groups, service principals, permissions, entitlements, and access-control models.
Implement role-based and attribute-based access controls, least-privilege permissions, separation of duties, and privileged-access procedures.
Configure secure access patterns for secrets, tokens, credentials, service principals, private endpoints, storage, and external systems.
Enable audit logging, lineage, system tables, tagging, data classification, row-level security, column masking, and compliance reporting.
Partner with security and governance teams on encrypti on, key management, network controls, data loss prevention, retention, auditability, and regulatory requirements.
Review access, monitor privileged activities, remediate policy violations, and support internal and external audits.
Cloud Infrastructure and Networking
Build and operate Databricks on AWS or Azure, including secure integration with cloud storage, identity, networking, encryption, and monitoring services.
On AWS, work with S3, IAM, KMS, VPC, PrivateLink, security groups, Route 53, CloudWatch, Secrets Manager, and related services.
On Azure, work with ADLS Gen2, Microsoft Entra ID, managed identities, Key Vault, virtual networks, private endpoints, network security groups, Azure Monitor, and related services.
Configure control-plane and data-plane connectivity, private networking, DNS, routing, firewall, proxy, and egress controls.
Integrate Databricks with cloud data lakes, APIs, databases, message platforms, and enterprise applications.
Support Kubernetes, Docker, EKS or AKS, and container-based platform services where required.
Contribute to capacity planning, disaster recovery, high availability, backup, restoration, and business-continuity exercises.
Infrastructure as Code and Automation
Build and maintain reusable Terraform modules using cloud and Databricks providers.
Automate workspace, network, Unity Catalog, storage, identity, compute-policy, cluster, job, permission, and monitoring configurations.
Manage Terraform state, workspaces, variables, modules, versioning, policy checks, drift detection, and controlled promotion across environments.
Use Python, shell scripting, Databricks CLI, REST APIs, and SDKs to automate administrative and operational tasks.
Implement self-service workspace, catalog, schema, access, and compute vending with appropriate approval and governance controls.
Maintain configuration standards and reduce manual administration through repeatable automation.
DevOps, CI/CD, and Release Engineering
Design and support CI/CD pipelines using GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps, or Harness.
Automate deployment of notebooks, jobs, workflows, libraries, policies, infrastructure, and platform configuration.
Support Databricks Asset Bundles, Git integration, artifact management, environment promotion, testing, approvals, and rollback.
Integrate security scanning, policy validation, infrastructure testing, and release evidence into delivery pipelines.
Enable engineering teams through templates, reusable pipelines, documentation, and self-service platform capabilities.
Partner with application, data-engineering, and DevOps teams to ensure deploy ment standards are consistent and supportable.
Reliability, Monitoring, Operations, and FinOps
Establish monitoring, alerting, dashboards, logs, metrics, traces, and health checks for Databricks and connected cloud services.
Use CloudWatch, Azure Monitor, Datadog, Splunk, New Relic, or similar platforms to monitor availability, compute utilization, failures, security events, and cost.
Define platform service-level indicators, service-level objectives, operational metrics, and error budgets.
Lead incident response, problem management, root-cause analysis, corrective actions, and post-incident reviews.
Manage vulnerability remediation, runtime patching, dependency updates, security exceptions, and platform lifecycle activities.
Optimize cluster sizing, autoscaling, pools, SQL warehouses, job concurrency, serverless usage, storage, and workload scheduling.
Implement budget controls, chargeback or showback tagging, utilization reporting, anomaly detection, and cost-optimization recommendations.
Participate in operational support rotations and maintain escalation paths with Databricks and cloud providers.
Coordinate platform upgrades, disaster-recovery tests, security reviews, and production-readiness assessments.
Collaboration and Technical Leadership
Partner with architecture, data engineering, security, cloud, network, governance, FinOps, and service-management teams.
Advise engineering teams on Databricks platform standards, secure patterns, deployment models, performance, and cost.
Conduct technical reviews and ensure solutions meet enterprise architecture and operational-support requirements.
Mentor platform engineers and administrators and lead knowledge-transfer sessions.
Communicate platform health, risks, dependencies, incidents, and improvement roadmaps to technical and business stakeholders.
Drive continuous improvement in automation, reliability, security, developer experience, and operational efficiency.
Required Qualifications
Typically 7-10 years of cloud, DevOps, Site Reliability Engineering, infrastructure, or platform-engineering experience.
At least 3 years of hands-on Databricks platform administration in an enterprise environment.
Strong experience administering Databricks workspaces, Unity Catalog, compute, cluster policies, jobs, SQL warehouses, permissions, and service principals.
Strong experience with AWS or Azure infrastructure, identity, storage, networking, encryption, monitoring, and private connectivity.
Strong proficiency with Terraform and infrastructure-as-code practices.
Experience with Python, shell script
By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively "Judge") to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.
Contact:
This job and many more are available through The Judge Group. Please apply with us today!
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: cxjudgpa
- Position Id: 1149944
- Posted 7 hours ago