Cloud Operations Manager

Hybrid in Philadelphia, PA, US • Posted 60+ days ago • Updated 4 days ago
Full Time
No Travel Required
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Cloud Operations Manager
  • Cloud Operations Lead
  • Cloud Infrastructure Manager
  • Cloud Infrastructure Lead
  • AWS Operations Manager
  • AWS Cloud Operations
  • Cloud Engineering Manager
  • Cloud Engineering Lead
  • Cloud Manager
  • AWS Manager
  • Cloud Infrastructure
  • AWS Cloud
  • AWS Architecture
  • AWS Infrastructure
  • Cloud Operations
  • Cloud Migration
  • Cloud Modernization
  • Site Reliability Engineering
  • SRE
  • DevOps
  • Infrastructure as Code
  • IaC
  • Terraform
  • CloudFormation
  • CI/CD
  • Jenkins
  • GitHub Actions
  • AWS CodePipeline
  • AWS EC2
  • AWS VPC
  • AWS IAM
  • S3
  • RDS
  • CloudWatch
  • EKS
  • ECS
  • Lambda
  • Systems Manager
  • Python
  • Bash
  • Cloud Security
  • AWS Security
  • IAM
  • Monitoring
  • Observability
  • Incident Management
  • Problem Management
  • Root Cause Analysis
  • RCA
  • High Availability
  • Scalability
  • Disaster Recovery
  • Business Continuity
  • SLA
  • SLO
  • SLI
  • Performance Optimization
  • Capacity Planning
  • FinOps
  • AWS Cost Optimization
  • Cloud Cost Management
  • Compliance
  • SOC2
  • HIPAA
  • GDPR
  • Vendor Management
  • Technical Leadership
  • Infrastructure Management
  • Production Operations
  • Enterprise Cloud.

Summary

Job Title: Cloud Operations Manager

Location: Philadelphia, PA
Experience: 12–18 Years
Employment Type: Full Time
Job Type: Permanent / Direct Hire

Job Summary

We are seeking an experienced Cloud Operations Manager to lead and optimize enterprise cloud infrastructure operations, reliability, automation, security, and service delivery.

The ideal candidate will have strong hands-on expertise in AWS Cloud, Cloud Operations, Site Reliability Engineering (SRE), Infrastructure as Code (IaC), Terraform, CloudFormation, CI/CD, monitoring, incident management, cloud security, and cost optimization.

This role will lead cross-functional technical teams including Cloud Engineers, SREs, Network Engineers, and Database Administrators, while partnering with technology, business, project, and vendor teams to ensure reliable, scalable, secure, and supportable production environments.

Required Qualifications

  • 12–18 years of overall IT infrastructure, cloud, or technology experience.

  • 5+ years of experience managing AWS cloud environments and enterprise cloud operations.

  • Proven experience leading and managing Cloud Engineering, SRE, Network, and Database teams.

  • Strong understanding of AWS cloud architecture, infrastructure, operations, reliability, and service management.

  • Strong experience with Infrastructure as Code (IaC) using Terraform and/or AWS CloudFormation.

  • Experience designing and managing CI/CD pipelines and DevOps automation.

  • Strong knowledge of cloud security, compliance, identity, access management, and data protection.

  • Experience with 24x7 production operations, monitoring, incident management, troubleshooting, and Root Cause Analysis (RCA).

  • Strong experience with cloud performance, scalability, availability, disaster recovery, and resilience.

  • Experience with cloud cost management, FinOps, and AWS cost optimization.

  • Strong leadership, strategic planning, communication, and stakeholder-management skills.

Key Responsibilities

<>Cloud Operations & Infrastructure
  • Lead day-to-day operations of enterprise AWS cloud infrastructure and production environments.

  • Ensure high availability, reliability, scalability, performance, and operational stability of cloud platforms.

  • Define and implement cloud operational standards, processes, procedures, and best practices.

  • Establish operational models supporting production maintainability, supportability, and service reliability.

  • Oversee infrastructure delivery for new applications, platforms, and software components.

  • Partner with architecture and engineering teams to ensure solutions are designed for operability, scalability, resiliency, and supportability.

<>Team Leadership
  • Lead, mentor, and develop teams of:

    • Cloud Engineers

    • Site Reliability Engineers (SRE)

    • Network Engineers

    • Database Administrators (DBAs)

    • DevOps / Infrastructure Engineers

  • Establish engineering and operational best practices.

  • Drive technical innovation, automation, collaboration, and operational excellence.

  • Manage team priorities, technical escalations, performance, and resource planning.

<>Infrastructure as Code & Automation
  • Drive adoption of Infrastructure as Code (IaC) across cloud infrastructure.

  • Design and implement automated infrastructure provisioning using Terraform and AWS CloudFormation.

  • Automate cloud operations, configuration management, deployment, monitoring, and remediation.

  • Develop reusable infrastructure modules, templates, and automation patterns.

  • Promote DevOps and CI/CD practices across cloud operations.

<>Monitoring, Reliability & Incident Management
  • Establish proactive cloud monitoring, observability, alerting, and operational health practices.

  • Ensure production systems meet defined availability, reliability, performance, and SLA/SLO requirements.

  • Lead major incident response, technical escalations, troubleshooting, and Root Cause Analysis (RCA).

  • Implement proactive problem management and preventative measures.

  • Drive Site Reliability Engineering (SRE) practices and continuous service improvement.

  • Develop and maintain operational dashboards, metrics, and reporting.

<>Security & Compliance
  • Implement and enforce enterprise cloud security and compliance standards.

  • Ensure AWS environments follow security best practices for identity, access, networking, data protection, and workload security.

  • Implement least privilege, IAM, encryption, logging, monitoring, vulnerability management, and security controls.

  • Support compliance requirements and regulatory standards such as:

    • SOC 2

    • HIPAA

    • GDPR

    • Other applicable industry and regulatory requirements

  • Identify cloud infrastructure risks and vulnerabilities and drive remediation.

<>Performance & Cost Optimization
  • Monitor AWS infrastructure performance and capacity.

  • Identify opportunities to improve availability, scalability, performance, and latency.

  • Drive AWS cost optimization and FinOps initiatives.

  • Analyze cloud consumption, utilization, and spending trends.

  • Recommend right-sizing, resource optimization, and architectural improvements without compromising reliability or performance.

<>Cloud Strategy & Transformation
  • Develop and execute enterprise cloud operations and infrastructure strategy.

  • Support cloud migration, modernization, scalability, and technology transformation initiatives.

  • Evaluate emerging AWS services and cloud technologies and recommend solutions that improve operational efficiency.

  • Establish cloud standards, reference architectures, operational patterns, and governance frameworks.

<>Vendor & Stakeholder Management
  • Partner with program managers, project managers, software engineering leaders, architects, business stakeholders, and third-party vendors.

  • Coordinate with AWS and technology vendors to resolve service issues and leverage new capabilities.

  • Support vendor evaluation, service management, and technical negotiations.

  • Communicate operational risks, service health, infrastructure priorities, and technology recommendations to leadership.

Must-Have Technical Skills

Cloud & Infrastructure

  • AWS

  • AWS Cloud Architecture

  • AWS Infrastructure

  • Cloud Operations

  • Cloud Infrastructure Management

  • Cloud Migration

  • Cloud Modernization

  • Enterprise Infrastructure

  • High Availability

  • Scalability

  • Disaster Recovery

  • Business Continuity

AWS Services

  • EC2

  • VPC

  • IAM

  • S3

  • RDS

  • CloudWatch

  • Lambda

  • EKS / ECS

  • Route 53

  • CloudFormation

  • AWS Organizations

  • AWS Systems Manager

Automation / IaC

  • Terraform

  • AWS CloudFormation

  • Infrastructure as Code (IaC)

  • Infrastructure Automation

  • Configuration Management

  • Python

  • Bash / Shell Scripting

DevOps / CI/CD

  • CI/CD

  • DevOps

  • Git

  • Jenkins

  • GitHub Actions

  • GitLab CI/CD

  • AWS CodePipeline / CodeBuild

  • Automated Deployment

SRE & Operations Skills

  • Site Reliability Engineering (SRE)

  • Production Operations

  • 24x7 Operations

  • Incident Management

  • Problem Management

  • Change Management

  • Root Cause Analysis (RCA)

  • Incident Response

  • Service Reliability

  • Observability

  • Monitoring & Alerting

  • Logging

  • Performance Management

  • Capacity Planning

  • SLA / SLO / SLI

  • Disaster Recovery

  • Business Continuity

  • Operational Excellence

Cloud Security & Compliance

  • AWS Security

  • Cloud Security

  • AWS IAM

  • Identity and Access Management

  • Least Privilege

  • Encryption

  • Data Protection

  • Vulnerability Management

  • Security Monitoring

  • Compliance

  • SOC 2

  • HIPAA

  • GDPR

  • Security Governance

  • Risk Management

FinOps & Cost Optimization

  • AWS Cost Optimization

  • Cloud Cost Management

  • FinOps

  • Resource Right-Sizing

  • Capacity Optimization

  • Cloud Spend Management

  • Cost Governance

  • Resource Utilization

  • AWS Cost Explorer

  • Cloud Financial Management

Leadership & Management Skills

  • Cloud Operations Leadership

  • Technical Team Management

  • Cross-Functional Leadership

  • Strategic Planning

  • Resource Planning

  • Vendor Management

  • Stakeholder Management

  • Executive Communication

  • Technical Mentoring

  • Operational Governance

  • Service Management

  • Continuous Improvement

Preferred Qualifications

  • AWS certifications such as:

    • AWS Certified Solutions Architect

    • AWS Certified SysOps Administrator

    • AWS Certified DevOps Engineer

  • Experience managing large-scale enterprise AWS environments.

  • Experience leading SRE / DevOps transformation initiatives.

  • Experience implementing cloud observability and monitoring platforms.

  • Experience with Kubernetes and containerized environments.

  • Experience with Terraform Enterprise / Terraform Cloud.

  • Experience with ITIL, Service Management, or related operational frameworks.

  • Experience managing cloud environments in highly regulated or enterprise environments.

Core Dice Search Keywords

Cloud Operations Manager, Cloud Operations Lead, Cloud Infrastructure Manager, Cloud Infrastructure Lead, AWS Operations Manager, AWS Cloud Operations, Cloud Engineering Manager, Cloud Engineering Lead, Cloud Manager, AWS Manager, Cloud Infrastructure, AWS Cloud, AWS Architecture, AWS Infrastructure, Cloud Operations, Cloud Migration, Cloud Modernization, Site Reliability Engineering, SRE, DevOps, Infrastructure as Code, IaC, Terraform, CloudFormation, CI/CD, Jenkins, GitHub Actions, AWS CodePipeline, AWS EC2, AWS VPC, AWS IAM, S3, RDS, CloudWatch, EKS, ECS, Lambda, Systems Manager, Python, Bash, Cloud Security, AWS Security, IAM, Monitoring, Observability, Incident Management, Problem Management, Root Cause Analysis, RCA, High Availability, Scalability, Disaster Recovery, Business Continuity, SLA, SLO, SLI, Performance Optimization, Capacity Planning, FinOps, AWS Cost Optimization, Cloud Cost Management, Compliance, SOC2, HIPAA, GDPR, Vendor Management, Technical Leadership, Infrastructure Management, Production Operations, Enterprise Cloud.

Education

Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent professional experience.

Location

Philadelphia, PA

Employment Type

Full Time / Permanent

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: mategr
  • Position Id: TAR-10
  • Posted 30+ days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Camden, New Jersey

Today

Easy Apply

Full-time

USD 140,000.00 - 150,000.00 per year

Hybrid in Trenton, New Jersey

Today

Easy Apply

Third Party, Contract

Depends on Experience

Hybrid in Trenton, New Jersey

4d ago

Easy Apply

Contract, Third Party

$50 - $55

Hybrid in Trenton, New Jersey

Today

Easy Apply

Third Party, Contract

Depends on Experience

Search all similar jobs