Senior Platform Engineer (Observability)

Evansville, IN, US • Posted 7 hours ago • Updated 7 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Partnership
  • Critical Thinking
  • Systems Analysis
  • Application Development
  • Finance
  • Presentations
  • Documentation
  • Process Improvement
  • Workflow
  • Reliability Engineering
  • Computer Science
  • Instrumentation
  • Grafana
  • AppDynamics
  • Data Collection
  • Dashboard
  • Communication
  • Problem Solving
  • Conflict Resolution
  • Management
  • Collaboration
  • Scripting
  • Bash
  • Windows PowerShell
  • Python
  • C
  • JavaScript
  • Agile
  • Performance Tuning
  • Capacity Management
  • Configuration Management
  • Continuous Improvement
  • Time Series
  • Ansible
  • Terraform
  • JSON
  • ServiceNow
  • Cloud Computing
  • Amazon Web Services
  • Microsoft Azure
  • Military
  • SAP BASIS
  • Law

Summary

We're seeking a Senior Monitoring Engineer to join a high-performing Monitoring Engineering team in a fast-paced finance technology organization. You'll design, develop, and maintain monitoring and observability solutions that keep core applications and infrastructure healthy and visible. In close partnership with application, platform, and development teams, you will implement alerting systems, dashboards, correlations, and automation-driving reliability, reducing MTTR, and elevating operational awareness.

Critical thinking, system analysis, and proactive troubleshooting are essential to success in this role.

Key Responsibilities

Design, Build, and Maintain Monitoring & Observability Solutions
  • Develop and maintain instrumentation, telemetry, and alerting for the Enterprise Monitoring Center using industry-leading tools, such as:
    • Grafana
    • OpsRamp
    • AppDynamics
    • Elastic Stack
    • BigPanda
    • AWS CloudWatch
    • Azure Monitor
  • Implement Observability best practices, ensuring comprehensive coverage of metrics, logs, and traces across critical systems.
  • Integrate and manage OpenTelemetry for distributed tracing and telemetry data collection, enabling end-to-end visibility of business-critical transactions.

Collaboration & Project Participation
  • Collaborate with application development teams to define and document observability requirements for each project or release.
  • Participate in complex initiatives, ensuring accurate and actionable monitoring and tracing are in place for every step of business-critical workflows.

Alerting & Escalation Process
  • Define and maintain standardized alert payloads per engineering guidelines, ensuring alerts are actionable.
  • Partner with Level 2 and Level 3 support teams to reflect process changes in monitoring dashboards.
  • Maintain and optimize thresholds, ensuring seamless escalations via BigPanda as the central alert hub.

Dashboard Creation & Maintenance
  • Create and maintain intuitive, actionable dashboards for the Enterprise Monitoring Center and other finance teams.
  • Ensure dashboards are effectively monitored by Level 1 teams, presenting clear, actionable data that reduces MTTR.

System Validation, Documentation & Automation
  • Develop and maintain automation scripts to enhance monitoring efficiency and improve team quality of life.
  • Proactively identify process improvements and learning opportunities; drive continuous improvement.

Automation & Quality-of-Life Improvements
  • Contribute to the automation of monitoring, alerting, and operational tasks to streamline workflows and improve overall system reliability.

Qualifications

Education

Bachelor's in Computer Science, IT, or related field.

Experience
  • Minimum 4 years in a technology organization, with 1 year hands-on engineering experience in monitoring or production operations.

Required Skills
  • Strong experience developing instrumentation and alerting for large, complex environments.
  • Expertise in 4 of the following: OpsRamp, Grafana, AppDynamics, Elastic Stack, InfluxDB, BigPanda, and other monitoring solutions.
  • Hands-on experience with Observability concepts and frameworks, including metrics, logs, and traces.
  • Working knowledge of OpenTelemetry for distributed tracing and telemetry data collection.
  • Experience with dashboard creation, alert management, and tool configuration.
  • Excellent verbal and written communication-able to present complex technical issues to both technical and non-technical stakeholders.
  • Strong problem-solving and troubleshooting in high-pressure environments.
  • Ability to prioritize and manage multiple tasks in a deadline-driven setting.
  • Proven collaboration with cross-functional teams in large, complex IT environments.
  • Experience with scripting (e.g., Bash, PowerShell) and proficiency in one programming language (e.g., Python, C family, JavaScript).
  • Experience designing and implementing scalable, reliable monitoring solutions.
  • Experience with agile software development methodologies
  • Familiar with problem diagnosis; performance tuning; capacity planning and configuration management across the stack via continuous improvement.

Preferred Qualifications
  • Experience querying, manipulating, and visualizing time-series data.
  • Familiarity with Infrastructure as Code tools (e.g., Ansible, Terraform).
  • Strong understanding of how to create actionable, digestible visualizations for Level 1 monitoring teams.
  • Working knowledge of REST APIs, JSON, and ServiceNow.
  • Experience with cloud monitoring-particularly AWS or Azure.

OneMain Holdings, Inc. is an Equal Employment Opportunity (EEO) employer. Qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship status, color, creed, culture, disability, ethnicity, gender, gender identity or expression, genetic information or history, marital status, military status, national origin, nationality, pregnancy, race, religion, sex, sexual orientation, socioeconomic status, transgender or on any other basis protected by law.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90970133
  • Position Id: eb320f570ebb1ecc1db4655d8bdaaf83
  • Posted 7 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote or Rhode Island

Today

Full-time

USD 83,430.00 per year

Remote or Rhode Island

Today

Full-time

USD 106,605.00 per year

Oklahoma

Today

Full-time

No location provided

Today

Easy Apply

Full-time, Contract, Third Party

$DOE

Search all similar jobs