Senior Observability & Monitoring Engineer

Remote • Posted 7 hours ago • Updated 7 hours ago
Full Time
Remote
Up to $100,000/yr
Fitment

Dice Job Match Score™

✨ Finding the perfect fit...

Job Details

Skills

  • Observability & Monitoring Engineer
  • Dynatrace
  • Moogsoft
  • Splunk
  • Cloud Computing
  • Google Cloud Platform
  • Google Cloud
  • Amazon Web Services
  • AWS
  • Python
  • Windows PowerShell
  • Scripting
  • Incident Management
  • Reliability Engineering
  • JIRA
  • ServiceNow

Summary

Senior Observability & Monitoring Engineer
Remote Role
Fulltime/ W2- Infinite Computer Solutions
About the Role
We are seeking an experienced Senior Observability & Monitoring Engineer to design, implement, and optimize enterprise monitoring and observability solutions supporting the migration of mission-critical applications from on-premises environments to AWS and Google Cloud Platform (Google Cloud Platform).
You will play a key role in improving system reliability by enabling proactive monitoring, intelligent alerting, rapid incident detection, and operational excellence.
Key Responsibilities
  • Design and implement enterprise-wide monitoring and observability solutions.
  • Build dashboards, KPIs, SLIs, and SLOs for applications, infrastructure, databases, APIs, and cloud services.
  • Configure intelligent monitoring, alerting, event correlation, and anomaly detection using Dynatrace, Splunk, and Moogsoft.
  • Develop synthetic monitoring and automated health checks for critical business services.
  • Partner with Cloud Architects, Developers, SREs, and Operations teams to ensure production readiness during cloud migrations.
  • Automate monitoring configurations and operational processes using scripting and AI-assisted tools.
Required Qualifications
  • 8+ years of experience in Monitoring, Observability, SRE, or Operations Engineering.
  • Strong hands-on experience with Dynatrace, Splunk, and Moogsoft.
  • Experience supporting cloud migrations to AWS and/or Google Cloud Platform.
  • Solid understanding of APM, distributed tracing, logging, metrics, and alert management.
  • Experience with Jira, ServiceNow, and incident management.
  • Scripting skills in Python, Bash, or PowerShell.
  • Strong knowledge of enterprise applications and distributed systems.
Preferred Qualifications
  • Experience in fintech, banking, or other regulated industries.
  • Knowledge of OpenTelemetry, Terraform, and Infrastructure as Code (IaC).
  • Experience with self-healing automation and cloud-native observability.
What Success Looks Like
  • Establish scalable monitoring standards for cloud migration initiatives.
  • Improve issue detection and reduce Mean Time to Resolution (MTTR).
  • Minimize alert fatigue through intelligent alert optimization.
  • Deliver reusable observability frameworks that enhance operational reliability.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10199915
  • Position Id: 9049931
  • Posted 7 hours ago
Contact the job poster
Shamim Ahmed

Shamim Ahmed

Sr Technical Recruiter @ Infinite Computer Solutions (ICS)
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Full-time

Remote or Arizona

Today

Full-time

depends on experience

Remote or Texas

Today

Full-time

depends on experience

Remote or Raleigh, North Carolina

Today

Full-time

depends on experience

Search all similar jobs