Position Summary
We are seeking a highly skilled Site Reliability Engineering (SRE) Lead to drive the reliability, scalability, performance, and operational excellence of enterprise cloud and data platforms. The ideal candidate will have hands-on experience with Azure DevOps, Microsoft Fabric, Semantic Models, and Kubernetes (TKE), along with a strong background in automation, observability, and cloud operations.
This role will lead reliability initiatives, implement modern DevOps practices, and ensure the availability and performance of critical analytics and application workloads.
Key Responsibilities
Site Reliability Engineering
- Define and implement SRE best practices, including SLIs, SLOs, and error budgets.
- Ensure high availability, scalability, and performance of cloud-native applications and data platforms.
- Lead incident management, root cause analysis (RCA), and problem management activities.
- Develop disaster recovery, backup, and business continuity strategies.
- Drive operational excellence through automation and continuous improvement.
Azure DevOps & Automation
- Design and maintain CI/CD pipelines using Azure DevOps.
- Automate infrastructure provisioning and application deployments using Infrastructure as Code (IaC).
- Implement DevSecOps practices and governance controls.
- Monitor deployment health and improve release reliability.
Microsoft Fabric Administration & Operations
- Manage and support Microsoft Fabric environments, including:
- Data Factory
- Data Engineering
- Lakehouse
- Data Warehouse
- Real-Time Analytics
- Power BI Workloads
- Ensure platform stability, governance, and performance optimization.
- Support Fabric deployment pipelines and workspace administration.
- Monitor capacity utilization and optimize resource consumption.
Semantic Model Management
- Support and optimize enterprise semantic models and Power BI datasets.
- Ensure reliable refresh schedules and data availability.
- Implement governance standards for shared semantic models.
- Troubleshoot model performance and query optimization issues.
Kubernetes (TKE) Platform Support
- Deploy, manage, and support Kubernetes clusters using Tanzu Kubernetes Engine (TKE).
- Monitor cluster health, workload performance, and resource utilization.
- Automate scaling, patching, and deployment processes.
- Troubleshoot containerized applications and infrastructure issues.
- Implement Kubernetes security and operational best practices.
Monitoring & Observability
- Implement and maintain monitoring solutions using:
- Azure Monitor
- Application Insights
- Log Analytics
- Prometheus
- Grafana
- Develop dashboards, alerts, and operational reporting.
- Identify and resolve performance bottlenecks proactively.
Leadership & Collaboration
- Lead and mentor a small team of SRE/DevOps engineers.
- Collaborate with Data Engineering, BI, Platform Engineering, and Application Development teams.
- Establish reliability standards and operational procedures.
- Participate in architecture reviews and technical decision-making.
Required Qualifications
Education
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
Experience
- 5+ years of experience in SRE, DevOps, Cloud Engineering, Platform Engineering, or Infrastructure Operations.
- Experience supporting enterprise cloud and data platforms.
- Experience leading technical initiatives and driving operational improvements.
Technical Skills
Azure & DevOps
- Azure DevOps Pipelines, Repos, and Release Management.
- CI/CD implementation and deployment automation.
- Infrastructure as Code (Terraform, Bicep, or ARM Templates).
- Azure cloud services and administration.
Microsoft Fabric
- Microsoft Fabric administration and operations.
- Fabric Workspaces, Capacity Management, and Deployment Pipelines.
- Experience with Lakehouse, Data Warehouse, and Data Engineering workloads.
Semantic Layer
- Power BI Semantic Models and Dataset Management.
- Performance tuning and governance.
- Experience supporting enterprise reporting environments.
Kubernetes (TKE)
- Tanzu Kubernetes Engine (TKE) administration.
- Kubernetes cluster operations and troubleshooting.
- Helm, Containerization, Networking, and Security fundamentals.
Monitoring & Automation
- Azure Monitor, Application Insights, Log Analytics.
- Prometheus and Grafana.
- PowerShell, Python, Bash, or similar scripting languages.