Job Title: Splunk Monitoring & Observability Engineer
Location: Remote
Employment Type: Full-time
Job Description:
Splunk Monitoring & Observability Engineer
Must Have Technical/Functional Skills
Role Overview
We are seeking an experienced Splunk Monitoring & Observability Engineer with strong hands-on expertise in Splunk monitoring. The role will focus on reviewing and enhancing the existing monitoring landscape, optimizing Splunk usage and alerting, identifying monitoring gaps, and proposing modern observability techniques to improve proactive detection, troubleshooting, and end-to-end service visibility.
Mandatory Skills
Strong hands-on experience with Splunk Enterprise and/or Splunk Cloud.
Expertise in SPL, dashboards, alerts, reports, and log analytics.
Experience in monitoring optimization, alert tuning, and alert rationalization.
Strong troubleshooting and Root Cause Analysis (RCA) skills.
Understanding of observability principles: Logs, Metrics, Traces, and Events.
Experience monitoring applications, APIs, databases, and infrastructure.
Knowledge of cloud monitoring across AWS, Azure, or Google Cloud Platform.
Automation/scripting experience using Python, Shell, or similar technologies.
Preferred Skills
Splunk ITSI, Splunk Observability Cloud, OpenTelemetry, Grafana, Prometheus, Dynatrace, AppDynamics, Datadog, Kubernetes, AWS CloudWatch, CI/CD, Ansible, and AIOps.
Expected Outcome
The candidate should act not only as a Splunk Monitoring Engineer, but as a Monitoring & Observability SME who can: Assess Identify Gaps Optimize Rationalize Propose Observability Improvements Implement Continuous Improvements.
Roles & Responsibilities
Design, develop, and enhance monitoring solutions using Splunk Enterprise/Splunk Cloud.
Develop and optimize SPL queries, dashboards, reports, alerts, and correlation rules.
Monitor applications, infrastructure, databases, APIs, middleware, and batch processes.
Review the current monitoring environment and identify gaps, duplicate/noisy alerts, ineffective thresholds, and performance issues.
Propose and implement monitoring optimization and alert rationalization initiatives.
Analyze incidents and production issues to support troubleshooting and root cause analysis.
Assess observability maturity across Logs, Metrics, Traces, and Events.
Propose modern observability techniques such as distributed tracing, service health monitoring, SLIs/SLOs, and proactive monitoring.
Enable end-to-end visibility across applications, infrastructure, cloud, and dependent services.
Identify opportunities for automation, monitoring-as-code, AIOps, and proactive/self-healing capabilities.
Collaborate with Application, Infrastructure, Cloud, SRE, DevOps, and Operations teams.