Data Platform Administration Architect
<>Location: Remote
Employment Type: Contract>
Position Overview
We are seeking an experienced Data Platform Administration Architect to lead the architecture, administration, governance, operational excellence, and continuous improvement of an enterprise data platform.
The ideal candidate will have deep hands-on expertise with Databricks and Azure Data Services, combined with strong experience in enterprise platform administration, production support, release and change management, incident response, ITSM, security, governance, and platform modernization.
This role will serve as a technical and operational leader, partnering closely with Data Engineering, Infrastructure, Security, Networking, Governance, and Business teams to ensure the data platform is secure, scalable, highly available, compliant, and operationally reliable.
Key Responsibilities
- Lead the architecture, administration, governance, and operational support of enterprise data platforms across production and non-production environments.
- Administer and support Databricks, Azure Data Factory (ADF), Azure Data Lake Storage (ADLS), Azure Synapse Analytics, Azure SQL, and related Azure services.
- Own platform availability, reliability, scalability, performance, security, and operational excellence.
- Define and maintain platform standards, environment strategies, governance frameworks, monitoring standards, and operational procedures.
- Manage platform provisioning, capacity planning, workload management, performance optimization, and cloud cost optimization.
- Establish and maintain comprehensive monitoring, observability, alerting, and platform health dashboards using Azure Monitor, Log Analytics, and related tools.
- Lead enterprise Change Management and Release Management processes, including CRs, SCRs, ECRs, deployment planning, rollback strategies, and post-release validation.
- Participate in Change Advisory Board (CAB) activities by providing technical assessments, risk analysis, and deployment readiness recommendations.
- Lead production support activities involving Incidents, Service Requests, Problems, and Major Incidents.
- Coordinate cross-functional teams during P1/P2/P3 production incidents and ensure timely service restoration.
- Conduct detailed Root Cause Analysis (RCA) and develop corrective and preventive action plans to minimize recurring incidents.
- Develop and maintain operational runbooks, knowledge articles, escalation procedures, disaster recovery plans, and support documentation.
- Establish and report operational KPIs, including platform availability, SLA compliance, release success rate, MTTR, incident trends, and platform health.
- Implement and govern CI/CD pipelines, deployment automation, and Infrastructure as Code (IaC) using Azure DevOps and related technologies.
- Ensure compliance with enterprise security standards, access governance requirements, audit controls, and SOX/regulatory requirements.
- Drive platform modernization, automation, operational maturity, and continuous improvement initiatives.
- Mentor platform engineers and administrators while providing technical leadership and architectural direction.
- Partner with vendors and strategic technology teams to evaluate and implement platform improvements.
Required Qualifications & Skills
- 10+ years of experience in enterprise data platforms, platform engineering, cloud infrastructure, platform administration, or data operations.
- Extensive hands-on experience administering Databricks and Azure data platforms in enterprise production environments.
- Strong expertise in:
- Databricks
- Azure Data Lake Storage (ADLS)
- Azure Synapse Analytics
- Azure SQL
- Azure Data Factory (ADF)
- Azure DevOps
- Azure IAM, RBAC, Service Principals, and Managed Identities
- Azure Monitor and Log Analytics
- Proven experience as a Data Platform Administrator, Databricks Administrator, Azure Platform Administrator, Platform Architect, or Platform Operations Lead.
- Strong understanding of Databricks environment provisioning, configuration, security, governance, workload management, and production support.
- Experience with capacity planning, performance tuning, platform optimization, and cloud cost management.
- Strong experience with enterprise monitoring, observability, alerting, and operational dashboards.
- Hands-on experience with ITIL-aligned Incident, Change, Release, Problem, Service Request, and Major Incident Management processes.
- Strong experience with ServiceNow or equivalent ITSM platforms.
- Proven experience leading enterprise release management, deployment execution, rollback planning, and post-release validation.
- Experience managing and resolving critical P1/P2/P3 production incidents.
- Strong experience preparing RCA reports and corrective action plans.
- Experience establishing support models, escalation procedures, on-call processes, and operational readiness reviews.
- Strong knowledge of CI/CD, Infrastructure as Code, deployment automation, and Azure DevOps best practices.
- Experience with cloud security, access governance, compliance, and audit controls.
- Experience supporting SOX-compliant, regulated, or audit-driven environments.
- Strong leadership, communication, stakeholder management, and problem-solving skills.
- Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field.