Mid-Level Observability Engineer

Irving, TX, US • Posted 1 day ago • Updated 5 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • High Availability
  • Performance Monitoring
  • Software Performance Management
  • Software Development
  • Lifecycle Management
  • Instrumentation
  • Microservices
  • Database
  • Dashboard
  • Root Cause Analysis
  • Artificial Intelligence
  • Performance Metrics
  • Software Release Life Cycle
  • Reporting
  • Service Level
  • Accountability
  • DevOps
  • Reliability Engineering
  • System Administration
  • Management
  • Software Security
  • Cloud Computing
  • Google Cloud
  • Google Cloud Platform
  • Orchestration
  • Kubernetes
  • Docker
  • Scripting
  • Scripting Language
  • Python
  • Bash
  • Windows PowerShell
  • API
  • Incident Management
  • IT Service Management
  • ServiceNow
  • JIRA
  • Communication
  • Terraform
  • Ansible
  • Dynatrace
  • Continuous Delivery
  • GitHub
  • GitLab
  • Continuous Integration
  • Jenkins

Summary

Job Description

Summary:

As our infrastructure and application landscape scales, ensuring high availability, performance, and reliability is critical. We are looking for a dedicated Observability Engineer to help us see clearly into our complex systems, predict issues before they impact customers, and empower our engineering teams with actionable insights.

As a Mid-Level Observability Engineer, you will be the driving force behind our application performance monitoring (APM) and infrastructure visibility. Acting as the resident Dynatrace subject matter expert, you will design, implement, and maintain monitoring solutions across our cloud-native and legacy environments. You will partner closely with DevOps, SRE, and software development teams to build a culture of proactive monitoring, define Service Level Objectives (SLOs), and automate alert remediation.

Duties and Responsibilities:
  • Dynatrace Administration: Own the deployment, configuration, and lifecycle management of Dynatrace OneAgent, ActiveGates, and integrations across all environments.
  • Instrumentation & Telemetry: Partner with development teams to instrument microservices, databases, and third-party applications to ensure full-stack visibility.
  • Dashboarding & Alerting: Create intuitive, role-based dashboards and configure intelligent, actionable alerts that reduce alert fatigue and accurately trigger on degraded user experiences.
  • Synthetic Monitoring: Design and maintain synthetic checks to simulate user behavior and monitor critical business transactions.
  • Root Cause Analysis: Leverage Dynatrace's Davis AI to assist incident response teams in rapidly identifying and resolving bottlenecks and outages.
  • Automation & CI/CD: Integrate observability into the deployment pipeline (e.g., using Dynatrace Keptn or standard CI/CD tools) to evaluate performance metrics during the build and release phases.
  • SLI/SLO Management: Help define, measure, and report on Service Level Indicators and Service Level Objectives to hold the engineering organization accountable for reliability.

Requirements:
  • Experience: 3-5 years of experience in DevOps, Site Reliability Engineering (SRE), or Systems Administration, with at least 2 years heavily focused on observability and monitoring.
  • Dynatrace Expertise: Strong, hands-on experience deploying and managing Dynatrace (OneAgent, ActiveGate, Synthetics, Application Security, and Log Monitoring).
  • Cloud & Containerization: Working knowledge of modern cloud environments (Google Cloud Platform) and container orchestration platforms (Kubernetes, Docker).
  • Scripting & Automation: Proficiency in at least one scripting language (Python, Bash, Go, or PowerShell) to automate administrative tasks and API interactions.
  • Incident Management: Familiarity with integrating monitoring tools with ITSM and alerting platforms (ServiceNow, Jira).
  • Communication: Strong ability to explain complex technical issues to both engineering teams and non-technical stakeholders.

Preferred Skills (Nice to Haves):
  • Current Dynatrace Certifications (e.g., Dynatrace Associate or Professional Certification).
  • Experience configuring Infrastructure as Code (IaC) tools like Terraform or Ansible to automate agent deployments.
  • Familiarity with OpenTelemetry standards and how to ingest OTel data into Dynatrace.
  • Understanding of CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins) and automated performance gates.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10507796
  • Position Id: 59cac80bfdd5551c5182c19458095720
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Coppell, Texas

Today

Easy Apply

Full-time

$55 - $65 per hour

Coppell, Texas

Today

Easy Apply

Contract

Remote or Hybrid in Dallas, Texas

29d ago

Easy Apply

Third Party, Contract

Depends on Experience

Plano, Texas

Today

Full-time

USD 131,427.00 - 191,500.00 per year

Search all similar jobs