Job Description :
Job Requirements Key Responsibilities Look for Local candidates to Atlanta it is 5 days onsite role. Lead and mentor a team of observability engineers supporting enterprise platforms and services.Define and execute the observability strategy, standards, and roadmap.Oversee monitoring, logging, alerting, tracing, and dashboarding solutions.Drive service reliability, incident response readiness, and operational excellence initiatives.Collaborate with application, infrastructure, cloud, and SRE teams to improve system health and performance.Establish KPIs, SLAs, and operational metrics to measure platform reliability and team effectiveness.Manage hiring, performance development, resource planning, and stakeholder communications.Ensure adoption of best practices for observability, automation, and proactive problem management.Champion the adoption of AI and Copilot-enabled workflows within the Observability organization.Evaluate and implement AI-driven monitoring, alert correlation, and incident management capabilities.Partner with engineering and platform teams to build intelligent operational dashboards and automated remediation solutions.QualificationsBachelor's degree in Computer Science, Engineering, or a related field.10+ years of experience in infrastructure, operations, SRE, platform engineering, or observability domains.3+ years of people management experience leading technical teams.Strong knowledge of observability platforms such as Splunk, Datadog, AppDynamics, Dynatrace, Grafana, Prometheus, OpenTelemetry, or similar tools.Experience working in cloud environments (Azure & Google Cloud Platform).Experience leveraging Microsoft Copilot, Generative AI, and AI-powered observability capabilities to improve operational efficiency, incident response, and engineering productivity.Knowledge of AI-assisted troubleshooting, anomaly detection, root cause analysis, and predictive monitoring solutions.Excellent communication, stakeholder management, and leadership skills.PreferredExperience leading globally distributed teams.Strong background in automation, DevOps, and reliability engineering practices.Familiarity with enterprise-scale monitoring and incident management processes.This role will be responsible for building a high-performing observability team that enables proactive detection, rapid troubleshooting, and improved reliability across business-critical services.Work Experience 10-15Years