Senior Observability Operations Engineer

Phoenix, AZ, US • Posted 1 day ago • Updated 1 day ago
Full Time
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Amazon Web Services
  • Ansible
  • Analytical Skill
  • Artificial Intelligence
  • DevOps
  • DNS
  • Continuous Integration
  • Continuous Delivery

Summary

Senior Observability Operations Engineer

Job Overview

We are seeking a highly skilled Senior Observability Operations Engineer to manage, optimize, and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions.

Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is highly desirable.

The role is responsible for ensuring high availability, scalability, operational excellence, and continuous improvement of enterprise monitoring and logging platforms supporting mission-critical applications.

Key Responsibilities

  • Administer and optimize enterprise observability platforms, including Dynatrace, Splunk, and OpenSearch/Elasticsearch.

  • Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.

  • Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.

  • Configure and administer Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, Real User Monitoring (RUM), Digital Experience Monitoring (DEM), Davis AI, and Application Performance Monitoring (APM).

  • Administer Splunk Enterprise, including Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.

  • Develop dashboards, alerts, reports, and executive operational metrics.

  • Support Linux-based infrastructure and Kubernetes environments, with Docker, OpenShift, or Rancher experience preferred.

  • Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.

  • Perform root cause analysis (RCA) for production incidents using observability platforms.

  • Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.

  • Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible.

  • Participate in incident, problem, change, and release management processes.

  • Drive platform upgrades, patching, security compliance, and operational governance.

  • Improve platform reliability through automation, self-healing, and AI-assisted operations.

Required Technical Skills

Observability Platforms

  • Dynatrace Administration

  • Splunk Enterprise Administration

  • OpenSearch Administration

  • Elasticsearch Administration

  • Grafana

  • Prometheus

  • Kibana

  • Jaeger

  • OpenTelemetry

  • Kafka — Preferred

Infrastructure

  • Linux Administration

  • Kubernetes

  • Docker

  • OpenShift or Rancher

  • Networking fundamentals, including TCP/IP, DNS, Load Balancers, and Firewalls

  • System Administration

Cloud & DevOps

Experience with one or more major cloud platforms:

  • AWS

  • Microsoft Azure

  • Google Cloud Platform (Google Cloud Platform)

Additional skills:

  • CI/CD Pipelines

  • Git

  • Terraform

  • Ansible

  • REST APIs

Scripting

  • Python

  • Bash/Shell Scripting

  • PowerShell — Preferred

AI/ML & Automation Skills — Preferred

  • Experience using Generative AI tools such as ChatGPT, GitHub Copilot, Amazon Q, Microsoft Copilot, or similar tools to improve operational efficiency.

  • Knowledge of AIOps platforms and AI-driven observability.

  • Experience with Dynatrace Davis AI for anomaly detection and root cause analysis.

  • Understanding of machine learning concepts for predictive monitoring and intelligent alerting.

  • Experience building AI-assisted operational runbooks and troubleshooting workflows.

  • Knowledge of Retrieval-Augmented Generation (RAG), vector databases, embeddings, and AI-powered knowledge search is a plus.

  • Experience integrating AI capabilities with observability platforms through APIs.

  • Familiarity with LLMs, prompt engineering, and AI-assisted automation.

  • Experience using Python with AI frameworks such as LangChain, LangGraph, OpenAI APIs, or similar technologies is desirable.

  • Exposure to AI-driven incident summarization, log analysis, and automated ticket enrichment.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or an equivalent combination of education and experience.

  • 6–10+ years of IT infrastructure or observability operations experience.

  • 4+ years of experience administering Dynatrace, Splunk, OpenSearch, and/or Elasticsearch.

  • Strong Linux system administration experience.

  • Experience supporting enterprise-scale production environments.

  • Strong troubleshooting and analytical skills.

  • Excellent communication and stakeholder management skills.

Preferred Certifications

  • Dynatrace Associate or Professional Certification

  • Splunk Enterprise Certified Administrator

  • Elastic Certified Engineer

  • Kubernetes CKA/CKAD

  • AWS, Azure, or Google Cloud Platform Certification

  • ITIL Foundation

  • AI/ML or Generative AI Certification — Preferred

Soft Skills

  • Strong ownership and accountability.

  • Excellent problem-solving and analytical skills.

  • Ability to work independently with minimal supervision.

  • Strong collaboration skills across cross-functional teams.

  • Continuous learning mindset.

  • Ability to thrive in fast-paced, production-critical environments.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91175551
  • Position Id: 9069814
  • Posted 1 day ago

Company Info

About Prama Innovations India Pvt. Ltd.

Prama is a technology company that builds AI, Data, Cloud, and Platform solutions for organizations that want to transform how they operate, decide, and serve their customers.

We engineer secure, scalable, and intelligent systems that move enterprises from reactive processes to predictive, automated decision-making.

About_Company_One
Contact the job poster
DV

Denish Vaghela

Recruiter @ Prama Innovations India Pvt. Ltd.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs