The ideal candidate will design and operate secure, highly available recovery platforms across on-premises, cloud, virtualized, and containerized environments, with a strong focus on ransomware resilience, immutable backups, cyber recovery vaults, recovery automation, and SRE best practices.
-
Design, engineer, and maintain highly available enterprise backup and recovery platforms using SRE principles.
-
Define and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery services.
-
Develop automation to reduce operational toil and improve platform reliability.
-
Perform Root Cause Analysis (RCA) and implement permanent corrective actions.
-
Improve platform reliability, scalability, performance, availability, and recoverability.
-
Build proactive monitoring, alerting, and observability for backup and cyber recovery platforms.
-
Participate in incident response, major incident management, and recovery operations.
-
Design and administer enterprise backup solutions across on-premises, cloud, SaaS, virtual machines, physical servers, databases, Kubernetes/OpenShift, NAS, object storage, and enterprise applications.
-
Engineer immutable backup architectures supporting ransomware resilience.
-
Optimize backup performance, retention, replication, encryption, and recovery objectives.
-
Implement policy-based backup automation and lifecycle management.
-
Ensure backup and recovery environments meet defined RPO and RTO requirements.
-
Design and implement air-gapped recovery vaults, Clean Rooms, IRE environments, and immutable storage architectures.
-
Develop secure recovery workflows for cyberattack and ransomware scenarios.
-
Automate malware scanning, recovery-point validation, and recovery readiness checks.
-
Design and test recovery orchestration for severe cyber disruption scenarios.
-
Work closely with Cyber Security teams to develop ransomware resilience strategies.
-
Develop Infrastructure as Code and Recovery as Code solutions.
-
Create automated recovery runbooks using Ansible, Terraform, Python, PowerShell, and GitHub Actions.
-
Automate recovery validation, compliance reporting, and evidence generation.
-
Implement monitoring for backup success rates, replication health, cyber vault health, recovery readiness, storage utilization, and infrastructure dependencies.
-
Build dashboards for operational teams and executive leadership.
-
Integrate backup and recovery platforms with Dynatrace, Grafana, Prometheus, Splunk, and other enterprise monitoring tools.
-
Plan and execute cyber recovery exercises, Clean Room validation, air-gap recovery testing, isolated recovery exercises, BMR testing, and Disaster Recovery testing.
-
Validate application recoverability against defined business RTO/RPO requirements.
-
Prepare executive-level reporting on cyber recovery readiness, resilience testing, risks, and remediation activities.
-
Experience working within financial services, banking, insurance, or another highly regulated industry.
-
Experience supporting GSIB cyber resiliency programs.
-
Understanding of regulatory expectations from organizations such as the Federal Reserve, OCC, and FFIEC.
-
Experience with chaos engineering and resilience testing.
-
Strong understanding of SRE reliability metrics and operational excellence practices.
-
Experience implementing AIOps, intelligent monitoring, or predictive analytics.
-
Excellent troubleshooting and root cause analysis skills.
-
Ability to lead cross-functional technical recovery initiatives.
-
Strong communication and executive presentation skills.
-
Proven ability to influence engineering standards and improve operational reliability.
-
Strong commitment to automation, continuous improvement, and resilience engineering.