Data Platform Administration Architect
Position Summary
The Data Platform Administration Architect is responsible for the architecture, administration, governance, operational excellence, and continuous improvement of the enterprise data platform. This role owns the stability, scalability, security, availability, and supportability of cloud-based data platforms, with a primary focus on Databricks, Azure Data Services, and enterprise platform operations.
The ideal candidate combines deep technical expertise with strong operational leadership and has extensive experience managing production environments, release processes, change governance, incident response, service management, and platform modernization initiatives. This individual will serve as the technical lead for platform administration while partnering with engineering, infrastructure, security, governance, and business teams to ensure reliable and compliant data platform operations.
Required Skills & Experience
- 10+ years of experience in enterprise data platforms, platform engineering, cloud infrastructure, platform administration, or data operations.
- Hands-on experience administering and supporting Databricks and Azure data platforms in enterprise production environments.
- Some experience in supply chain domain is good to have
- Strong expertise with:
- Databricks
- Azure Data Lake Storage (ADLS)
- Azure Synapse Analytics
- Azure SQL
- Azure Data Factory (ADF)
- Azure DevOps
- Azure IAM (RBAC, Service Principals, Managed Identities)
- Azure Monitor and Log Analytics
- Proven experience as a Data Platform Administrator, Databricks Administrator, Azure Platform Administrator, Platform Architect, or Platform Operations Lead, not solely a data engineer or application developer.
- Deep understanding of standing up, configuring, securing, governing, and supporting production and non-production Databricks environments.
- Experience managing platform provisioning, environment strategy, capacity planning, workload management, performance tuning, cost optimization, and operational support.
- Strong experience implementing and managing platform monitoring, observability, alerting, and operational dashboards.
- Proven experience managing enterprise production support processes, including:
- Incidents (INC)
- Service Requests (RITM)
- Change Requests (CR)
- Standard Changes (SCR)
- Emergency Changes (ECR)
- Problem Management
- Hands-on experience coordinating and leading enterprise release management activities, including release planning, deployment execution, rollback planning, deployment validation, and post-release support.
- Strong knowledge of ITIL-aligned operational processes, including:
- Incident Management
- Change Management
- Release Management
- Problem Management
- Service Request Management
- Knowledge Management
- Major Incident Management
- Experience resolving and leading response efforts for critical production incidents (P1/P2/P3), including executive communications and stakeholder coordination.
- Demonstrated experience creating and presenting comprehensive Root Cause Analysis (RCA) reports and corrective action plans following production incidents.
- Experience establishing support operating models, on-call rotations, escalation procedures, and operational readiness reviews.
- Strong experience with ServiceNow or equivalent ITSM platforms for managing incidents, changes, requests, approvals, releases, and audit requirements.
- Experience developing operational metrics, service-level reporting, trend analysis, and platform health dashboards.
- Experience designing scalable platform controls and operational processes to support enterprise data volumes and growth.
- Strong experience with CI/CD automation, deployment pipelines, Infrastructure as Code (IaC), and Azure DevOps best practices.
- Experience with cloud governance, platform security, access management, and compliance controls.
- Experience supporting SOX-compliant, regulated, or audit-driven environments.
- Ability to partner effectively with infrastructure, security, networking, application, and business teams to drive platform reliability and operational excellence.
- Proven leadership experience managing platform, operations, or engineering teams.
- Strong communication, stakeholder management, leadership, and problem-solving skills.
- Bachelor's degree in Computer Science, Information Systems, Engineering, or related field.
Preferred Qualifications
- Databricks Certified Data Engineer, Databricks Platform Administrator, or related Databricks certifications.
- Microsoft Certified: Azure Administrator Associate.
- Microsoft Certified: Azure Solutions Architect Expert.
- ITIL Foundation or ITIL 4 Certification.
- Experience leading CAB meetings, release governance forums, or enterprise change management programs.
- Experience supporting large-scale enterprise Databricks and Azure data platforms across multiple regions and business units.
- Experience with observability platforms such as Azure Monitor, Datadog, Splunk, Grafana, or similar tools.
- Experience implementing enterprise platform governance, FinOps, security, and compliance programs.
- Experience leading cloud modernization, platform transformation, and operational maturity initiatives.
- Strong understanding of data governance, access governance, Unity Catalog, data security, and regulatory compliance requirements.
- Vendor management and strategic technology partnership experience.
Key Responsibilities
- Lead the administration, architecture, governance, and operational support of enterprise data platforms.
- Manage and support Databricks, Azure Data Factory, Azure Data Lake Storage (ADLS), Azure Synapse, Azure SQL, and related Azure services across production and non-production environments.
- Own platform reliability, availability, performance, scalability, security, and operational excellence.
- Define and maintain platform standards, environment strategies, monitoring frameworks, operational procedures, and governance controls.
- Oversee Change Management and Release Management processes, including CRs, SCRs, ECRs, deployment planning, release scheduling, rollback strategies, and post-release validation.
- Participate in CAB (Change Advisory Board) activities and provide technical assessment, risk evaluation, and deployment readiness recommendations.
- Lead production support activities, including resolution of Incidents (INC), Service Requests (RITM), Problem Records, and Major Incident Management processes.
- Coordinate cross-functional teams during critical production incidents and ensure timely restoration of services.
- Perform Root Cause Analysis (RCA) investigations and drive preventative and corrective actions to reduce recurring issues.
- Develop and maintain operational runbooks, knowledge articles, support documentation, escalation procedures, and disaster recovery plans.
- Establish and report on operational KPIs including platform uptime, SLA compliance, release success rates, Mean Time to Resolution (MTTR), incident trends, and platform health metrics.
- Implement and govern CI/CD processes, deployment automation, and infrastructure-as-code best practices.
- Ensure compliance with enterprise security standards, access governance policies, audit requirements, and SOX controls.
- Drive platform modernization initiatives, automation opportunities, cost optimization, and continuous service improvement efforts.
- Mentor platform engineers and administrators while providing technical leadership and architectural direction.