8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles.
You’ve operated observability infrastructure at serious scale. You know what breaks at 10x and you
design for it.
You have a strong bias toward simplicity. You’ve seen over-engineered observability stacks collapse
under their own weight and you build accordingly.
Deep hands-on experience with a significant subset of: Prometheus, Thanos, VictoriaMetrics, Grafana,
Loki, Tempo, OpenTelemetry, ClickHouse, Elastic.
Strong engineering fundamentals, proficient in Python, Go, or similar; comfortable owning complex
systems end to end.
Experience with Kubernetes at scale; familiarity with GPU infrastructure or HPC environments (Slurm)
is a strong plus.
You can architect systems, write the code, review others’ work, and explain the tradeoffs clearly, all in
the same week.
Infrastructure-as-Code is default, not optional (Terraform, Ansible, or equivalent).
You influence without authority. Teams want your opinion because it makes their work better.
Preferred
Experience with high-volume streaming pipelines for observability data (Kafka, Vector, Fluent Bit, etc.).
Background in AI/ML infrastructure observability: GPU utilisation, training job visibility, inference
latency.
Prior experience defining observability strategy at an organisation level.