Cloud Senior Monitoring and Observability Engineer

Remote • Posted 9 hours ago • Updated 9 hours ago
Full Time
Remote
USD $107,900.00 - 195,050.00 per year
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • ISS
  • Normalization
  • Performance Analysis
  • Cloud Computing
  • Storage
  • Middleware
  • Virtual Machines
  • Meta-data Management
  • Onboarding
  • Workflow
  • Incident Management
  • Root Cause Analysis
  • Leadership
  • Technical Writing
  • SolarWinds
  • Dynatrace
  • New Relic
  • Splunk
  • Microsoft Windows
  • Linux
  • Dashboard
  • Reporting
  • Ansible
  • Scripting
  • Continuous Integration
  • Continuous Delivery
  • IT Service Management
  • ServiceNow
  • Project Management
  • Employment Authorization
  • SEC
  • Security Clearance
  • Management
  • Software Performance Management
  • Log Management
  • Network
  • Performance Monitoring
  • Database
  • Kubernetes
  • Virtualization
  • API
  • FISMA
  • FedRAMP
  • Amazon Web Services
  • Microsoft Azure
  • Red Hat Linux
  • Terraform
  • ITIL
  • Recruiting
  • Market Analysis
  • Law

Summary

The Senior Monitoring and Observability Engineer will support the Leidos SEC ISS2 contract by engineering, operating, and continuously improving enterprise monitoring and observability capabilities across hybrid infrastructure, cloud, and container platforms.

This hands-on role is responsible for monitoring coverage, platform integration, agent deployment, tagging and normalization, dashboards, alerting, logs, APM, synthetic monitoring, automation, and operational integrations. Datadog is the primary enterprise observability platform used in the environment. Strong Datadog experience is preferred, however, candidates with substantial experience engineering and operating other enterprise monitoring or observability platforms will be considered where they demonstrate strong transferable monitoring expertise and the ability to rapidly develop proficiency with new technologies.

The engineer partners with Operations and engineering teams to improve visibility, alert quality, incident detection, troubleshooting, performance analysis, and operational reliability across the enterprise.

Primary Responsibilities

In this Role you will:
  • Engineer, operate, maintain, and continuously improve the enterprise monitoring and observability platform, including dashboards, monitors, metrics, logs, APM, synthetic monitoring, tagging, integrations, and related capabilities.
  • Assess monitoring coverage across enterprise systems and applications, identify visibility gaps, and coordinate onboarding or remediation with the appropriate technical teams.
  • Maintain monitoring coverage across Windows, Linux, cloud, OpenShift/Kubernetes, virtualized, database, network, storage, middleware, and application environments.
  • Support monitoring and observability for Red Hat OpenShift, Kubernetes, OpenShift Virtualization, and virtual machine workloads running on OpenShift.
  • Configure and troubleshoot monitoring agents, integrations, collectors, APIs, and related platform components.
  • Build and maintain consistent tagging, metadata, dashboards, alerts, service health views, and operational reporting.
  • Automate monitoring deployment, configuration, tagging, onboarding, upgrades, and integrations using Ansible, APIs, scripting, CI/CD, infrastructure-as-code, or similar technologies.
  • Develop and maintain integrations between observability platforms, ServiceNow, notification systems, on-call workflows, and other enterprise operational systems.
  • Use monitoring and observability data to troubleshoot complex performance and availability issues, support incident response and root-cause analysis, and recommend technical remediation.
  • Correlate infrastructure, application, platform, and dependency telemetry to identify service degradation and recurring technical issues.
  • Partner with Operations and engineering teams to improve monitoring coverage, alert quality, service visibility, incident detection, escalation, and operational response.
  • Analyze telemetry and historical trends to identify capacity risks, recurring issues, monitoring gaps, and opportunities for improvement.
  • Develop actionable performance, availability, capacity, and monitoring coverage reporting for technical and leadership stakeholders.
  • Maintain monitoring standards, technical documentation, configuration guidance, and operational procedures.

Basic Qualifications
  • BS degree and 8-12 years of prior relevant experience, or Master's degree with 6-10 years of prior relevant experience. Additional relevant experience may be considered in lieu of degree requirements where permitted by contract.
  • Strong hands-on experience engineering and operating enterprise monitoring or observability platforms.
  • Strong Datadog experience is preferred; however, substantial experience with ScienceLogic SL1, SolarWinds, Dynatrace, New Relic, Splunk Observability, LogicMonitor, PrometheGrafana, or comparable enterprise platforms will be considered based on demonstrated monitoring and observability engineering expertise.
  • Demonstrated ability to apply monitoring and observability engineering principles across technologies and rapidly develop proficiency with new platforms.
  • Production experience monitoring Windows and Linux infrastructure and Kubernetes or Red Hat OpenShift environments.
  • Experience deploying, configuring, upgrading, and troubleshooting monitoring agents, integrations, dashboards, alerts, tagging, and operational reporting.
  • Experience automating monitoring deployment or administration using Ansible, APIs, scripting, CI/CD pipelines, infrastructure-as-code, or similar technologies.
  • Experience integrating monitoring or observability platforms with ITSM systems such as ServiceNow.
  • Strong troubleshooting and dependency-analysis skills across infrastructure, applications, networks, platforms, and services.
  • Ability to analyze technical telemetry, identify monitoring or performance gaps, and translate findings into actionable recommendations.
  • Ability to communicate technical findings and recommendations to technical teams, project leadership, and customer stakeholders.
  • Must meet applicable contract citizenship and work authorization requirements and be able to obtain and maintain SEC Public Trust or other required clearance.

Preferred Qualifications
  • Direct experience engineering or administering Datadog in a large enterprise environment.
  • Experience with application performance monitoring, distributed tracing, or OpenTelemetry.
  • Experience with Datadog APM, Log Management, Synthetic Monitoring, RUM, Network Performance Monitoring, Database Monitoring, or related capabilities.
  • Experience with Red Hat OpenShift Virtualization, CNV, KubeVirt, or related Kubernetes-based virtualization technologies.
  • Experience monitoring Microsoft Azure or AWS environments.
  • Experience with Terraform, monitoring-as-code, API-driven deployment, or related infrastructure-as-code approaches.
  • Experience supporting federal agency IT environments governed by FISMA, FedRAMP, NIST, or related security requirements.
  • Relevant technical certifications such as Datadog, AWS, Microsoft Azure, Red Hat OpenShift, Terraform, or ITIL are preferred.

If you're looking for comfort, keep scrolling. At Leidos, we outthink, outbuild, and outpace the status quo - because the mission demands it. We're not hiring followers. We're recruiting the ones who disrupt, provoke, and refuse to fail. Step 10 is ancient history. We're already at step 30 - and moving faster than anyone else dares.

Original Posting:
August 24, 2026

For U.S. Positions: While subject to change based on business needs, Leidos reasonably anticipates that this job requisition will remain open for at least 3 days with an anticipated close date of no earlier than 3 days after the original posting date as listed above.

Pay Range:
Pay Range $107,900.00 - $195,050.00

The Leidos pay range for this job level is a general guideline only and not a guarantee of compensation or salary. Additional factors considered in extending an offer include (but are not limited to) responsibilities of the job, education, experience, knowledge, skills, and abilities, as well as internal equity, alignment with market data, applicable bargaining agreement (if any), or other law.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: SCNCAPI2
  • Position Id: 9b9419831084da066324392ab47d2d25
  • Posted 9 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

6d ago

Easy Apply

Full-time

Up to $100,000

Remote or Rhode Island

Today

Full-time

USD 83,430.00 per year

Remote

Today

Full-time

USD 86,800.00 - 198,000.00 per year

Remote

Today

Full-time

Search all similar jobs