Senior/Lead Site Reliability Engineer - Observability

Remote • Posted 5 hours ago • Updated 5 hours ago
Full Time
Remote
$100,000 - $130,000/yr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Splunk Enterprise
  • Splunk Cloud
  • Elasticsearch
  • ELK
  • Kibana
  • Prometheus
  • Grafana
  • Grafana Tempo
  • Open Telemetry
  • Distributed Tracing
  • Kafka
  • Terraform
  • Kubernetes
  • Docker
  • Linux
  • Python
  • Go
  • Ruby
  • Bash
  • AWS
  • Ansible
  • Consul
  • Site Reliability Engineering
  • Platform Engineering
  • DevOps
  • FedRAMP
  • Continuous Delivery
  • Continuous Integration
  • Dashboard
  • Good Clinical Practice
  • Google Cloud Platform
  • Analytics
  • Reliability Engineering
  • SPL
  • Servers
  • Splunk
  • IaaS
  • Apache Kafka
  • Cloud Computing
  • Configuration Management
  • Instrumentation
  • Microsoft Azure
  • Amazon Web Services

Summary

Job Role: Senior/Lead Site Reliability Engineer - Observability

Location: Remote

Job Description:

Must Have Technical/Functional Skills

  • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.
  • Hands-on experience administering Splunk Enterprise or Splunk Cloud.
  • Strong knowledge of Splunk SPL.
  • Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, Open Telemetry, and Kafka.
  • Experience implementing metrics, logs, and traces as part of a modern observability strategy.
  • Experience with Terraform and Infrastructure as Code.
  • Programming experience in Python, Go, Ruby, or Bash.
  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
  • Experience supporting FedRAMP or regulated environments.

Technology Stack

Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, Open Telemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.

Roles & Responsibilities:

  • Design, deploy, and operate enterprise observability platforms.
  • Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, Search Head Clusters, Heavy Forwarders, and Deployment Servers.
  • Deploy and operate large-scale Elasticsearch clusters for log analytics and search.
  • Design, deploy, and support distributed tracing platforms using Grafana Tempo and Open Telemetry.
  • Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.
  • Scale Prometheus, Grafana, Kafka, Tempo, and Open Telemetry-based monitoring solutions.
  • Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.
  • Automate infrastructure using Terraform and configuration management tools.

Nice to have skills:

  • Splunk certification.
  • Experience with Kubernetes, AWS/Azure/Google Cloud Platform, Ansible, Consul, CI/CD pipelines, and service mesh technologies.
  • Experience supporting FedRAMP or regulated environments.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10217521
  • Position Id: 9067542
  • Posted 5 hours ago
Contact the job poster
RB

Rajat Bhoyar

Recruiter @ Ztek Consulting
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Easy Apply

Full-time

Depends on Experience

Remote

Today

Full-time

USD 141,800.00 - 195,000.00 per year

Remote

2d ago

Easy Apply

Third Party, Contract

$63 - $65

Remote

Today

Full-time

Search all similar jobs