Job Title: Cloud Operations Manager
Location: Philadelphia, PA
Experience: 12–18 Years
Employment Type: Full Time
Job Type: Permanent / Direct Hire
Job Summary
We are seeking an experienced Cloud Operations Manager to lead and optimize enterprise cloud infrastructure operations, reliability, automation, security, and service delivery.
The ideal candidate will have strong hands-on expertise in AWS Cloud, Cloud Operations, Site Reliability Engineering (SRE), Infrastructure as Code (IaC), Terraform, CloudFormation, CI/CD, monitoring, incident management, cloud security, and cost optimization.
This role will lead cross-functional technical teams including Cloud Engineers, SREs, Network Engineers, and Database Administrators, while partnering with technology, business, project, and vendor teams to ensure reliable, scalable, secure, and supportable production environments.
Required Qualifications
12–18 years of overall IT infrastructure, cloud, or technology experience.
5+ years of experience managing AWS cloud environments and enterprise cloud operations.
Proven experience leading and managing Cloud Engineering, SRE, Network, and Database teams.
Strong understanding of AWS cloud architecture, infrastructure, operations, reliability, and service management.
Strong experience with Infrastructure as Code (IaC) using Terraform and/or AWS CloudFormation.
Experience designing and managing CI/CD pipelines and DevOps automation.
Strong knowledge of cloud security, compliance, identity, access management, and data protection.
Experience with 24x7 production operations, monitoring, incident management, troubleshooting, and Root Cause Analysis (RCA).
Strong experience with cloud performance, scalability, availability, disaster recovery, and resilience.
Experience with cloud cost management, FinOps, and AWS cost optimization.
Strong leadership, strategic planning, communication, and stakeholder-management skills.
Key Responsibilities
<>Cloud Operations & Infrastructure>
Lead day-to-day operations of enterprise AWS cloud infrastructure and production environments.
Ensure high availability, reliability, scalability, performance, and operational stability of cloud platforms.
Define and implement cloud operational standards, processes, procedures, and best practices.
Establish operational models supporting production maintainability, supportability, and service reliability.
Oversee infrastructure delivery for new applications, platforms, and software components.
Partner with architecture and engineering teams to ensure solutions are designed for operability, scalability, resiliency, and supportability.
<>Team Leadership>
Lead, mentor, and develop teams of:
Cloud Engineers
Site Reliability Engineers (SRE)
Network Engineers
Database Administrators (DBAs)
DevOps / Infrastructure Engineers
Establish engineering and operational best practices.
Drive technical innovation, automation, collaboration, and operational excellence.
Manage team priorities, technical escalations, performance, and resource planning.
<>Infrastructure as Code & Automation>
Drive adoption of Infrastructure as Code (IaC) across cloud infrastructure.
Design and implement automated infrastructure provisioning using Terraform and AWS CloudFormation.
Automate cloud operations, configuration management, deployment, monitoring, and remediation.
Develop reusable infrastructure modules, templates, and automation patterns.
Promote DevOps and CI/CD practices across cloud operations.
<>Monitoring, Reliability & Incident Management>
Establish proactive cloud monitoring, observability, alerting, and operational health practices.
Ensure production systems meet defined availability, reliability, performance, and SLA/SLO requirements.
Lead major incident response, technical escalations, troubleshooting, and Root Cause Analysis (RCA).
Implement proactive problem management and preventative measures.
Drive Site Reliability Engineering (SRE) practices and continuous service improvement.
Develop and maintain operational dashboards, metrics, and reporting.
<>Security & Compliance>
Implement and enforce enterprise cloud security and compliance standards.
Ensure AWS environments follow security best practices for identity, access, networking, data protection, and workload security.
Implement least privilege, IAM, encryption, logging, monitoring, vulnerability management, and security controls.
Support compliance requirements and regulatory standards such as:
Identify cloud infrastructure risks and vulnerabilities and drive remediation.
<>Performance & Cost Optimization>
Monitor AWS infrastructure performance and capacity.
Identify opportunities to improve availability, scalability, performance, and latency.
Drive AWS cost optimization and FinOps initiatives.
Analyze cloud consumption, utilization, and spending trends.
Recommend right-sizing, resource optimization, and architectural improvements without compromising reliability or performance.
<>Cloud Strategy & Transformation>
Develop and execute enterprise cloud operations and infrastructure strategy.
Support cloud migration, modernization, scalability, and technology transformation initiatives.
Evaluate emerging AWS services and cloud technologies and recommend solutions that improve operational efficiency.
Establish cloud standards, reference architectures, operational patterns, and governance frameworks.
<>Vendor & Stakeholder Management>
Partner with program managers, project managers, software engineering leaders, architects, business stakeholders, and third-party vendors.
Coordinate with AWS and technology vendors to resolve service issues and leverage new capabilities.
Support vendor evaluation, service management, and technical negotiations.
Communicate operational risks, service health, infrastructure priorities, and technology recommendations to leadership.
Must-Have Technical Skills
Cloud & Infrastructure
AWS Services
EC2
VPC
IAM
S3
RDS
CloudWatch
Lambda
EKS / ECS
Route 53
CloudFormation
AWS Organizations
AWS Systems Manager
Automation / IaC
DevOps / CI/CD
SRE & Operations Skills
Cloud Security & Compliance
FinOps & Cost Optimization
Leadership & Management Skills
Cloud Operations Leadership
Technical Team Management
Cross-Functional Leadership
Strategic Planning
Resource Planning
Vendor Management
Stakeholder Management
Executive Communication
Technical Mentoring
Operational Governance
Service Management
Continuous Improvement
Preferred Qualifications
AWS certifications such as:
AWS Certified Solutions Architect
AWS Certified SysOps Administrator
AWS Certified DevOps Engineer
Experience managing large-scale enterprise AWS environments.
Experience leading SRE / DevOps transformation initiatives.
Experience implementing cloud observability and monitoring platforms.
Experience with Kubernetes and containerized environments.
Experience with Terraform Enterprise / Terraform Cloud.
Experience with ITIL, Service Management, or related operational frameworks.
Experience managing cloud environments in highly regulated or enterprise environments.
Core Dice Search Keywords
Cloud Operations Manager, Cloud Operations Lead, Cloud Infrastructure Manager, Cloud Infrastructure Lead, AWS Operations Manager, AWS Cloud Operations, Cloud Engineering Manager, Cloud Engineering Lead, Cloud Manager, AWS Manager, Cloud Infrastructure, AWS Cloud, AWS Architecture, AWS Infrastructure, Cloud Operations, Cloud Migration, Cloud Modernization, Site Reliability Engineering, SRE, DevOps, Infrastructure as Code, IaC, Terraform, CloudFormation, CI/CD, Jenkins, GitHub Actions, AWS CodePipeline, AWS EC2, AWS VPC, AWS IAM, S3, RDS, CloudWatch, EKS, ECS, Lambda, Systems Manager, Python, Bash, Cloud Security, AWS Security, IAM, Monitoring, Observability, Incident Management, Problem Management, Root Cause Analysis, RCA, High Availability, Scalability, Disaster Recovery, Business Continuity, SLA, SLO, SLI, Performance Optimization, Capacity Planning, FinOps, AWS Cost Optimization, Cloud Cost Management, Compliance, SOC2, HIPAA, GDPR, Vendor Management, Technical Leadership, Infrastructure Management, Production Operations, Enterprise Cloud.
Education
Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent professional experience.
Location
Philadelphia, PA
Employment Type
Full Time / Permanent