![]()
Job Description:
Position : Dev Ops Engineer
Location : Chicago, IL / Tempe, AZ / New Jersey
Pay Range: $75/hr - $80/hr on W2
Duration: 12 months contract, with highly possibilities or extensions or conversion
Position Summary
We are seeking a highly skilled Senior Site Reliability Engineer (SRE) with extensive expertise in enterprise backup engineering, cyber recovery, and platform resiliency. In this role, you will develop and maintain highly available, secure, and automated recovery systems to protect critical organizational services from operational failures, cyber threats such as ransomware, and other cyber risks. The ideal candidate is proficient in applying traditional SRE principles-automation, observability, reliability engineering, and resilience-along with experience designing and managing enterprise backup platforms, immutable storage solutions, air-gapped cyber vaults, isolated recovery environments, and recovery orchestration processes. You will work closely with Infrastructure, Cyber Security, Cloud Engineering, Application Development, and Disaster Recovery teams to ensure critical services are resilient, continuously validated, and always recoverable.
Key Responsibilities
- Engineer and sustain highly available, resilient enterprise platforms following SRE principles.
- Define, monitor, and improve Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery services.
- Develop automation solutions to minimize manual efforts and enhance system reliability.
- Conduct root cause analyses (RCA) and implement lasting corrective actions.
- Enhance platform scalability, performance, recoverability, and overall reliability.
- Establish proactive monitoring, alerting, and observability for backup and cyber recovery systems.
- Participate actively in incident response and major recovery operations.
- Design, implement, and oversee backup and recovery architectures across on-premises, cloud, and SaaS environments.
- Create immutable backup architectures to support ransomware resilience.
- Develop backup strategies for virtual environments, physical servers, databases, Kubernetes/OpenShift workloads, cloud-native applications, NAS/Object Storage, and enterprise applications.
- Optimize backup performance parameters including retention, replication, encryption, and recovery objectives.
- Implement policy-driven automation for backups and data lifecycle management to ensure compliance with RPO and RTO requirements.
- Design and deploy enterprise cyber recovery solutions such as air-gapped vaults, clean rooms, isolated recovery environments, and immutable storage architectures.
- Develop secure recovery workflows and automated malware scanning and validation processes following cyberattack scenarios.
- Test and validate recovery orchestration for severe cyber events; support recovery point validation and promotion workflows.
- Collaborate with Cyber Security teams on ransomware resilience and cyber recovery strategies.
- Develop Infrastructure as Code (IaC) and Recovery as Code automation, including building automated recovery runbooks with tools such as scripting languages, Terraform, and CI/CD pipelines.
- Automate recovery validation, reporting, and compliance documentation.
- Implement comprehensive observability and monitoring for backup success, replication health, recovery readiness, storage utilization, vault health, and infrastructure dependencies.
- Build dashboards for operational insights and integrate with enterprise observability platforms.
- Lead cyber resilience testing exercises, including cyber recovery drills, clean room validation, air-gapped recovery tests, full isolated recovery environment exercises, bare-metal recovery, and disaster recovery tests.
- Produce executive reports on recovery readiness and testing outcomes.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
- At least 7 years of experience in Backup Engineering, Infrastructure Engineering, or Site Reliability Engineering.
- Minimum of 5 years designing and implementing enterprise backup solutions.
- At least 3 years supporting cyber recovery architectures.
- Proven experience applying SRE principles within complex enterprise infrastructure environments.
- Strong understanding of distributed systems and high availability architectural design.
- Preferred Qualifications
- Prior experience working in highly regulated industries such as financial services.
- Knowledgeable about cyber resiliency programs aligned with industry regulators.
- Familiarity with regulatory standards from agencies such as the Federal Reserve, OCC, or FFIEC.
- Experience conducting resilience testing and chaos engineering exercises.
- Understanding of SRE tooling, reliability metrics, and performance monitoring frameworks.
- Experience implementing AI-assisted operations (AIOps) and predictive analytics to enhance reliability.
- Demonstrated leadership in cross-functional recovery efforts with strong communication and presentation skills.
- Ability to influence engineering standards and promote operational excellence through automation and reliability initiatives.
Dexian stands at the forefront of Talent + Technology solutions with a presence spanning more than 70 locations worldwide and a team exceeding 10,000 professionals. As one of the largest technology and professional staffing companies and one of the largest minority-owned staffing companies in the United States, Dexian combines over 30 years of industry expertise with cutting-edge technologies to deliver comprehensive global services and support.
Dexian connects the right talent and the right technology with the right organizations to deliver trajectory-changing results that help everyone achieve their ambitions and goals. To learn more, please visit .
Dexian is an Equal Opportunity Employer that recruits and hires qualified candidates without regard to race, religion, sex, sexual orientation, gender identity, age, national origin, ancestry, citizenship, disability, or veteran status.