Job Title: Cloud Operations Lead
Location: Remote (Occasional travel)
Job type: Contract
Key Responsibilities:
First-Line Hands-On Troubleshooting
· Act as the primary technical escalation point and hands-on operational responder for infrastructure, network, database, and cloud workload issues across Google Cloud Platform (primary) / AWS.
· Troubleshoot GitOps deployment failures, synchronization issues, and application rollouts using Argo CD and GitLab CI/CD.
· Perform real-time diagnostic analysis using observability tools, cloud-native logs, and metrics to identify and resolve root causes.
· Execute standard operational procedures, automate repetitive maintenance tasks, and apply immediate hotfixes or operational patches to restore service availability.
· This is a heavy troubleshooting and project-based role requiring active participation in an on-call rotation for production support.
Incident Management:
· Lead major incident management (MIM) processes, acting as the primary incident commander during outages, failed deployments, or high-severity events.
· Coordinate cross-functional technical teams to drive rapid resolution and minimize service impact (MTTR).
· Manage incident communication across stakeholders, leadership, and external partners.
· Own on-call rotations, incident escalation paths, and bridge management
Problem Management:
· Facilitate Post-Incident Reviews (PIRs) and root cause analysis (RCA) to uncover underlying systemic weaknesses.
· Track, prioritize, and manage problem records to prevent recurring pipeline, deployment, and infrastructure issues.
· Partner with SRE, DevOps, and Development teams to prioritize stability fixes, architectural remediations, and technical debt elimination.
Service Management & Operations Leadership
· Maintain operational SLAs, SLOs, and OLAs; continuously monitor and report on system uptime and operational metrics.
· Manage ITSM practices including Change Management, Event Management, Release Governance, and Capacity Planning.
· Lead, mentor, and build a team of cloud support engineers, fostering a continuous-improvement mindset around automated CI/CD and GitOps practices.
· Manage vendor relationships and cloud provider support agreements (Google Cloud Platform/AWS support tickets).
· Maintain strong security awareness across infrastructure and deployment practices, flagging and addressing risks proactively.
Qualifications & Skills:
Must Have:
· Kubernetes: Strong hands-on experience (deployments, troubleshooting, scaling, managed services like GKE/EKS).
· Argo CD: Expert-level, hands-on experience (application syncing, rollbacks, status monitoring, GitOps workflows).
· Database Knowledge: Solid working knowledge of databases (administration, performance troubleshooting, query-level debugging).
· Google Cloud Platform: Strong hands-on cloud experience (compute, networking, IAM, monitoring).
· Security Awareness: Working knowledge of cloud security best practices, IAM policies, and vulnerability awareness.
Technical Knowledge:
· GitOps & Deployment Tools: Deep hands-on experience with Argo CD and GitLab (CI/CD pipelines, runner management, repository management).
· OS & Networking: Deep knowledge of Linux administration, networking protocols (TCP/IP, DNS, VPNs, Firewalls), and cloud networking (VPCs, route tables, load balancers).
· Monitoring & Observability: Hands-on experience with tools such as Datadog, CloudWatch, Google Cloud Monitoring, Prometheus, Grafana, Splunk, or New Relic.
· Automation & IaC: Proficiency with scripting languages (Python, Bash) and Infrastructure as Code (Terraform, Ansible).
· Containerization: Strong practical experience with Docker, Kubernetes, and managed container services (GKE, EKS).
Process & Management Experience:
· ITSM Frameworks: Strong practical knowledge of ITIL guidelines (ITIL v4 certification is a plus).
· Leadership: 3+ years managing, leading, or mentoring cloud operations, support, or SRE teams.
· Troubleshooting: Proven ability to troubleshoot complex, distributed multi-tier architectures and automated CI/CD build/release failures under pressure.
· Communication: Excellent written and verbal communication skills for stakeholder updates and technical documentation.
· Availability: Willingness and ability to participate in on-call rotations, including off-hours incident response.
Preferred Requirements
· Bachelor's degree in Computer Science, Information Technology, or equivalent experience.
· Professional Cloud Certification (Google Cloud Associate Cloud Engineer / Professional Cloud Architect, or AWS equivalent).
· Kubernetes or GitOps certifications (CKA, CKAD, GitOps Certified Associate).
· Experience with incident management tools such as PagerDuty, Opsgenie, ServiceNow, or Jira Service Management.
Nice to Have:
· FinOps skills — familiarity with cloud cost governance and reporting.
· Demonstrated cost-saving experience — track record of identifying and executing cloud cost optimization initiatives.