Job Title: Senior Monitoring and Observability Engineer
Location: Washington, DC (Remote)
Type: Contract
Compensation: $63.00 - $75.00 per hour on W2
Security Clearance: Must be able to obtain a Public Trust clearance
Responsibilities
- Engineer, operate, maintain, and continuously improve the enterprise monitoring and observability platform, including dashboards, monitors, metrics, logs, APM, synthetic monitoring, tagging, integrations, and related capabilities.
- Assess monitoring coverage across enterprise systems and applications, identify visibility gaps, and coordinate onboarding or remediation with the appropriate technical teams.
- Maintain monitoring coverage across Windows, Linux, cloud, OpenShift/Kubernetes, virtualized, database, network, storage, middleware, and application environments.
- Support monitoring and observability for Red Hat OpenShift, Kubernetes, OpenShift Virtualization, and virtual machine workloads running on OpenShift.
- Configure and troubleshoot monitoring agents, integrations, collectors, APIs, and related platform components.
- Build and maintain consistent tagging, metadata, dashboards, alerts, service health views, and operational reporting.
- Automate monitoring deployment, configuration, tagging, onboarding, upgrades, and integrations using Ansible, APIs, scripting, CI/CD, infrastructure-as-code, or similar technologies.
- Develop and maintain integrations between observability platforms, ServiceNow, notification systems, on-call workflows, and other enterprise operational systems.
- Use monitoring and observability data to troubleshoot complex performance and availability issues, support incident response and root-cause analysis, and recommend technical remediation.
- Correlate infrastructure, application, platform, and dependency telemetry to identify service degradation and recurring technical issues.
- Partner with Operations and engineering teams to improve monitoring coverage, alert quality, service visibility, incident detection, escalation, and operational response.
- Analyze telemetry and historical trends to identify capacity risks, recurring issues, monitoring gaps, and opportunities for improvement.
- Develop actionable performance, availability, capacity, and monitoring coverage reporting for technical and leadership stakeholders.
- Maintain monitoring standards, technical documentation, configuration guidance, and operational procedures.
Requirements
- BS degree and 8-12 years of prior relevant experience, or Master's degree with 6-10 years of prior relevant experience.
- Strong hands-on experience engineering and operating enterprise monitoring or observability platforms.
- Strong Datadog experience is preferred; however, substantial experience with ScienceLogic SL1, SolarWinds, Dynatrace, New Relic, Splunk Observability, LogicMonitor, PrometheGrafana, or comparable enterprise platforms will be considered based on demonstrated monitoring and observability engineering expertise.
- Demonstrated ability to apply monitoring and observability engineering principles across technologies and rapidly develop proficiency with new platforms.
- Production experience monitoring Windows and Linux infrastructure and Kubernetes or Red Hat OpenShift environments.
- Experience deploying, configuring, upgrading, and troubleshooting monitoring agents, integrations, dashboards, alerts, tagging, and operational reporting.
- Experience automating monitoring deployment or administration using Ansible, APIs, scripting, CI/CD pipelines, infrastructure-as-code, or similar technologies.
- Experience integrating monitoring or observability platforms with ITSM systems such as ServiceNow.
- Strong troubleshooting and dependency-analysis skills across infrastructure, applications, networks, platforms, and services.
- Ability to analyze technical telemetry, identify monitoring or performance gaps, and translate findings into actionable recommendations.
- Ability to communicate technical findings and recommendations to technical teams, project leadership, and customer stakeholders.
- Must meet applicable contract citizenship and work authorization requirements and be able to obtain and maintain SEC Public Trust or other required clearance.
- Relevant technical certifications such as Datadog, AWS, Microsoft Azure, Red Hat OpenShift, Terraform, or ITIL are preferred.
System One, and its subsidiaries including Joulé and Mountain Ltd., are leaders in delivering outsourced services and workforce solutions across North America. We help clients get work done more efficiently and economically, without compromising quality. System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.
System One is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, age, national origin, disability, family care or medical leave status, genetic information, veteran status, marital status, or any other characteristic protected by applicable federal, state, or local law.
#M-M2
#LI-RF1
Ref: #851-Rockville-S1