Description
We are looking for an experienced Site Reliability Engineer to strengthen observability and operational excellence across a Microsoft Azure environment.. This Long-term Contract position will work closely with DevOps and engineering teams to create reliable monitoring practices, improve system insight, and support stable production operations. The role is ideal for someone who can translate reliability goals into practical monitoring solutions, actionable alerts, and measurable service performance improvements.
Responsibilities:
Create and evolve an observability framework that supports Azure-based platforms as well as integrated third-party services.
Develop meaningful dashboards, alerting rules, log analysis workflows, and distributed tracing capabilities using platforms such as Datadog or Dynatrace.
Utilize Azure monitoring services, including Azure Monitor, Log Analytics, Application Insights, Managed Prometheus, and Azure Managed Grafana, to expand operational visibility.
Evaluate system behavior across applications, infrastructure, and dependent services to improve performance tracking and service health awareness.
Collaborate with DevOps and software engineering teams to strengthen reliability practices, accelerate issue detection, and improve response to production incidents.
Analyze current monitoring coverage to uncover blind spots, reduce unnecessary alert volume, and support faster root-cause identification.
Support infrastructure and application reliability initiatives within cloud-native environments, including Kubernetes-based deployments and Azure-hosted services.
Contribute to implementation efforts tied to CI/CD workflows and infrastructure automation to ensure observability is consistently embedded in delivery processes.
Requirements
At least 5 years of hands-on experience working within Microsoft Azure environments.
Strong practical knowledge of Datadog, Dynatrace, or a comparable enterprise observability platform.
Proven experience with Azure Monitor, Log Analytics, Application Insights, Prometheus, and Grafana.
Background working with GitHub Actions, Terraform or other Infrastructure as Code tools, and Kubernetes.
Solid understanding of monitoring concepts such as metrics, logs, traces, alerting strategies, and production reliability.
Ability to independently lead observability efforts from solution design through implementation and ongoing refinement.
Experience supporting production systems, troubleshooting operational issues, and applying site reliability engineering practices such as incident response, SLIs, and SLOs.
Bachelor's degree in Computer Science, Information Technology, or a related discipline.
Technology Doesn't Change the World, People Do.
Robert Half is the world's first and largest specialized talent solutions firm that connects highly qualified job seekers to opportunities at great companies. We offer contract, temporary and permanent placement solutions for finance and accounting, technology, marketing and creative, legal, and administrative and customer support roles.
Robert Half works to put you in the best position to succeed. We provide access to top jobs, competitive compensation and benefits, and free online training. Stay on top of every opportunity - whenever you choose - even on the go. Download the Robert Half app and get 1-tap apply, notifications of AI-matched jobs, and much more.
All applicants applying for U.S. job openings must be legally authorized to work in the United States. Benefits are available to contract/temporary professionals, including medical, vision, dental, and life and disability insurance. Hired contract/temporary professionals are also eligible to enroll in our company 401(k) plan. Visit roberthalf.gobenefits.net for more information.
2025 Robert Half. An Equal Opportunity Employer. M/F/Disability/Veterans. By clicking "Apply Now," you're agreeing to Robert Half's Terms of Use and Privacy Notice.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: rhalfint
- Position Id: 03340-0013519435
- Posted 2 hours ago