Observability data collection and automation to help lead transformational initiatives within IT operations, encompassing development as well. As a crucial figure in this role, you will participate/help with various technology domain groups and cross functional teams on unified observability gap analysis and solutioning (automation and manual fixes)
Incorporate GenAI tooling and agentic capabilities to strengthen reliability outcomes across monitoring/alerting, rapid incident response, change management/testing, and DevOps/deployment processes.
Experience building agentic workflows using LLMs, tool-calling, function-calling, multi-agent orchestration, and event-driven automation.
Experience with Agent-to-Agent communication, AI agent federation, and enterprise AI control-plane concepts.
Experience implementing AI control-plane governance, including policy-based execution, approval workflows, audit trails, guardrails, and risk-based remediation controls.
Expertise in Observability as a service, Dashboard as a services, monitoring as a services and alert as a service in all technology domains (application, infrastructure, database, security, middleware, network etc.,) Telemetry data collection using Dynatrace APM, SolarWinds, CISCO Switches, F5, Databases, Open-Source tools (Prometheus and Grafana), Log Aggregations (Kibana or Splunk) and AIOPS Tools.
Practical experience implementing Golden Signals (latency, traffic, errors, saturation) using related telemetry sources.
Configure application performance monitoring (APM), infrastructure monitoring, synthetic monitoring, RUM, and log monitoring.
Integrate Dynatrace with CI/CD pipelines, alerting tools, ITSM systems, and incident automation frameworks.
Tune alert thresholds, baselines, and AI-driven anomaly detection to reduce noise and improve actionable insights.