We are seeking a highly experienced Sr Site Reliability Engineer – Storage Platforms to design, implement, and support Software Defined Storage (SDS) and Kubernetes platforms in a private cloud environment. This role focuses on scalability, resilience, automation, and performance using Infrastructure-as-Code and GitOps practices.
This is a deeply technical role requiring expert-level understanding of Software Defined Storage, Kubernetes, and extensive working knowledge on Linux Operating systems. You will also collaborate with platform and SRE teams to maintain secure, performant, and multitenant-isolated services that serve high-throughput, mission-critical applications.
Key Responsibilities
• Design, implement, and operate large-scale Software Defined Storage architectures across private and public cloud regions within ITIL methodology.
• Deploy and support enterprise storage platforms (Pure Storage, HPE, NetApp) and SDS solutions (Ceph, Longhorn).
• Build self-service storage workflows for Kubernetes CSI and OpenStack consumers (VM and Baremetal).
• Develop Infrastructure-as-Code using Ansible, Terraform, Helm and Git, with Python/Bash automation.
• Implement CI/CD pipelines for infrastructure updates, patching, upgrades, testing, and rollback.
• Build observability, alerting, and auto-remediation using GitOps and tools such as Prometheus, Loki, and Grafana.
• Architect and maintain high availability, disaster recovery, and scale-out infrastructure.
• Develop and review high-level and low-level design documents for storage infrastructure
• Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux systems.
• Participate in on-call rotations, incident response, and root cause analysis.
• Collaborate globally on change management, documentation, and operational best
practices.