SRE Azure Cloud

Hybrid in Pleasanton, CA, US • Posted 11 hours ago • Updated 11 hours ago
Contract W2
Contract Corp To Corp
6 Months
No Travel Required
Hybrid
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Grafana
  • Docker
  • kubernetes
  • Azure
  • powershell
  • python
  • PostgreSQL
  • prometheus
  • MySQL
  • Helm charts
  • Bash
  • Azure DevOps Pipelines
  • Cosmos DB
  • Azure Kubernetes Service (AKS)
  • SRE Azure Cloud
  • Monitoring & Observability
  • CI/CD & DevOps

Summary

Position: SRE Azure Cloud-12+ Year exp required

Onsite in Pleasanton CA - F2F Interview required

Please share local profiles as there might be F2F round as well.

JD for SRE onsite

Ideal Candidate: A hands-on Senior SRE with expertise in Stage/Production deployments, Azure-based microservices, observability (Grafana, Prometheus, Graylog, Azure Monitor), troubleshooting using Application Insights, database operations, and customer escalation management, focused on maintaining highly reliable and available production services.

 

Position Summary

We are seeking a highly skilled Senior Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational stability of customer-facing applications and services. This role is focused on Stage/Production deployments, Monitoring & Observability, Troubleshooting, Incident Management, and Customer Escalation support across Azure-based microservices environments.

The ideal candidate will possess strong experience in cloud operations, microservices, production support, deployment automation, observability platforms, and database troubleshooting, with a proven ability to rapidly diagnose and resolve complex production issues.

Key Responsibilities

Production & Stage Operations

•            Manage and support Stage and Production environments.

•            Execute application, infrastructure, configuration, and database deployments.

•            Validate releases, perform health checks, and coordinate rollback activities.

•            Support change management and production readiness reviews.

Monitoring & Observability

•            Build and maintain dashboards, alerts, and monitoring solutions.

•            Monitor application, infrastructure, and database health using logs, metrics, traces, and telemetry.

•            Improve observability coverage and reduce alert noise.

•            Proactively identify reliability and performance issues before customer impact.

Troubleshooting & Incident Response

•            Troubleshoot software, infrastructure, configuration, deployment, and database-related issues.

•            Lead incident response activities and production recovery efforts.

•            Perform root cause analysis (RCA) and implement preventive actions.

•            Develop operational runbooks and troubleshooting documentation.

Customer Escalation Management

•            Investigate and resolve customer-reported production issues.

•            Act as a technical lead during high-priority incidents.

•            Partner with Engineering, Product, and Customer Support teams to drive issue resolution.

•            Provide timely communication and status updates during major incidents.

Required Technical Skills

Cloud & Infrastructure

•            Microsoft Azure

•            Azure Kubernetes Service (AKS)

•            Azure Virtual Machines

•            App Services

•            Azure Storage

•            Azure Networking

•            Application Gateway

•            Azure Key Vault

Microservices & Containerization

•            Kubernetes

•            Docker

•            Helm Charts

•            Microservices Architecture

•            REST APIs

•            Event-Driven Architecture

•            Distributed Systems Troubleshooting

CI/CD & DevOps

•            Azure DevOps Pipelines

•            Bitbucket

•            Git

•            Helm-based Deployments

•            CI/CD Release Management

•            Deployment Automation

Monitoring & Observability

•            Grafana

•            Prometheus

•            Graylog

•            Azure Monitor

•            Application Insights

•            Log Analytics

•            Alerting & Dashboard Management

•            Distributed Tracing

•            SLI/SLO Monitoring

Troubleshooting Expertise

•            Application Performance Issues

•            Production Incident Management

•            Configuration & Environment Issues

•            Deployment Failures & Rollbacks

•            Kubernetes & Container Troubleshooting

•            Network & Connectivity Issues

•            Root Cause Analysis (RCA)

Databases

•            Azure SQL / SQL Server

•            PostgreSQL / MySQL

•            Cosmos DB

•            Redis

•            Query Performance Tuning

•            Database Monitoring

•            Backup & Recovery

Automation & Scripting

•            PowerShell

•            Python

•            Bash

Preferred Experience

•            Supporting enterprise SaaS applications in Production environments.

•            Azure-based microservices platforms running on AKS.

•            Customer-facing production support and escalation management.

•            24x7 on-call and incident response environments.

•            Site Reliability Engineering (SRE) best practices including SLIs, SLOs, MTTR, and service availability management.

Key Competencies

•            Strong troubleshooting and analytical skills.

•            Production support and incident management expertise.

•            Customer-first mindset.

•            Excellent communication and stakeholder management.

•            Ability to perform effectively during critical outages and high-severity incidents.

•            Continuous improvement and automation mindset.

Thanks & Regards 

Aryan Chaudhary 

Ex/106 /

Sr. Technical Recruiter 

DMS Vision Inc

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91133942
  • Position Id: 3131-35330-1790870900
  • Posted 11 hours ago

Company Info

About DMS Vision Inc.

At DMS Vision, our main goal is to be an integral part of our customer’s success. With ambition to be a Global Premier Provider of innovative, value-based technology solutions, our team has the drive and determination to do whatever it takes to meet the needs of our clients. Through our services, we strive to save our client’s money, time and hassle in every way possible.

At DMS Vision, we provide IT Staffing, Software Development, Cybersecurity and IoT Development and Services.

Whenever we take on a new project, we take extensive measures to learn all that we can take care about our client’s business. This allows us to better understand their goals and become familiar with the company’s philosophy.

We combine the insights obtained from this unique perspective with our professional strategic processes to develop a detailed plan for achieving the client’s ultimate vision of accomplishment. Our collaborative approach to problem solving and our technological expertise enables us to tackle even the most complex of our customer’s problems.

Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs