Minimum Qualifications
• Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
• Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
• Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
• Terraform or OpenTofu proficiency.
• Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
• Strong automation skills in Python, Bash, or Go.
• Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
• CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
• Proven ability to troubleshoot complex distributed systems, largely self-directed.
Preferred Qualifications
• GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar.
• NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
• Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
• Distributed tracing and OpenTelemetry instrumentation across services.
• Progressive delivery: canary and blue/green rollouts with automated rollback.
• Chaos or fault-injection testing, game days, and disaster-recovery drills.
• Multi-cloud networking, unified storage abstractions, and disaster recovery.
• FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
• Establishing an SRE function where one did not previously exist.
Thanks & Regards,
Narendra Kunware