Role: DevOps / SRE Cloud Lead
Location: 100% Remote
Key Responsibilities
First-Line Hands-On Troubleshooting
• Act as the primary technical escalation point and hands-on operational responder for infrastructure, network, database, and cloud workload issues across Google Cloud Platform (primary) / AWS.
• Troubleshoot GitOps deployment failures, synchronization issues, and application rollouts using Argo CD and GitLab CI/CD.
• Perform real-time diagnostic analysis using observability tools, cloud-native logs, and metrics to identify and resolve root causes.
• Execute standard operational procedures, automate repetitive maintenance tasks, and apply immediate hotfixes or operational patches to restore service availability.
• This is a heavy troubleshooting and project-based role requiring active participation in an on-call rotation for production support.
Incident Management
• Lead major incident management (MIM) processes, acting as the primary incident commander during outages, failed deployments, or high-severity events.
• Coordinate cross-functional technical teams to drive rapid resolution and minimize service impact (MTTR).
• Manage incident communication across stakeholders, leadership, and external partners.
• Own on-call rotations, incident escalation paths, and bridge management.
Problem Management
• Facilitate Post-Incident Reviews (PIRs) and root cause analysis (RCA) to uncover underlying systemic weaknesses.
• Track, prioritize, and manage problem records to prevent recurring pipeline, deployment, and infrastructure issues.
• Partner with SRE, DevOps, and Development teams to prioritize stability fixes, architectural remediations, and technical debt elimination.
Service Management & Operations Leadership
• Maintain operational SLAs, SLOs, and OLAs; continuously monitor and report on system uptime and operational metrics.
• Manage ITSM practices including Change Management, Event Management, Release Governance, and Capacity Planning.
• Lead, mentor, and build a team of cloud support engineers, fostering a continuous-improvement mindset around automated CI/CD and GitOps practices.
• Manage vendor relationships and cloud provider support agreements (Google Cloud Platform/AWS support tickets).
• Maintain strong security awareness across infrastructure and deployment practices, flagging and addressing risks proactively.
Qualifications & Skills
Must Have
• Kubernetes: Strong hands-on experience (deployments, troubleshooting, scaling, managed services like GKE/EKS).
• Argo CD: Expert-level, hands-on experience (application syncing, rollbacks, status monitoring, GitOps workflows).
• Database Knowledge: Solid working knowledge of databases (administration, performance troubleshooting, query-level debugging).
• Google Cloud Platform: Strong hands-on cloud experience (compute, networking, IAM, monitoring).
• Security Awareness: Working knowledge of cloud security best practices, IAM policies, and vulnerability awareness.
Technical Knowledge
• GitOps & Deployment Tools: Deep hands-on experience with Argo CD and GitLab (CI/CD pipelines, runner management, repository management).
• OS & Networking: Deep knowledge of Linux administration, networking protocols (TCP/IP, DNS, VPNs, Firewalls), and cloud networking (VPCs, route tables, load balancers).
• Monitoring & Observability: Hands-on experience with tools such as Datadog, CloudWatch, Google Cloud Monitoring, Prometheus, Grafana, Splunk, or New Relic.
• Automation & IaC: Proficiency with scripting languages (Python, Bash) and Infrastructure as Code (Terraform, Ansible).
• Containerization: Strong practical experience with Docker, Kubernetes, and managed container services (GKE, EKS).
Process & Management Experience
• ITSM Frameworks: Strong practical knowledge of ITIL guidelines (ITIL v4 certification is a plus).
• Leadership: 3+ years managing, leading, or mentoring cloud operations, support, or SRE teams.
• Troubleshooting: Proven ability to troubleshoot complex, distributed multi-tier architectures and automated CI/CD build/release failures under pressure.
• Communication: Excellent written and verbal communication skills for stakeholder updates and technical documentation.
• Availability: Willingness and ability to participate in on-call rotations, including off-hours incident response.
Preferred Requirements
• Bachelor's degree in Computer Science, Information Technology, or equivalent experience.
• Professional Cloud Certification (Google Cloud Associate Cloud Engineer / Professional Cloud Architect, or AWS equivalent).
• Kubernetes or GitOps certifications (CKA, CKAD, GitOps Certified Associate).
• Experience with incident management tools such as PagerDuty, Opsgenie, ServiceNow, or Jira Service Management.
Nice to Have
• FinOps skills — familiarity with cloud cost governance and reporting.
• Demonstrated cost-saving experience — track record of identifying and executing cloud cost optimization initiatives.