SRE Engineer

Remote • Posted 55 minutes ago • Updated 55 minutes ago
Contract Corp To Corp
Contract W2
Contract Independent
12 Months
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • SRE
  • AI
  • API
  • Infrastructure
  • MS Azure
  • DevOps
  • Terraform

Summary

Hi,

We are looking for a SRE Engineer @ Remote. Kindly send me your updated profile and I am looking forward to work with you in this position.

Job Title: SRE Engineer

Location: Remote (but need to travel based on request)

Duration: Long term contract

Reporting Line: Cloud and Infrastructure Lead

Key Skills:

  • Full stack SRE
  • Coding
  • AI platform built
  • AI partners
  • API gateways
  • Infrastructure
  • Code base-quality
  • High caliber
  • Manage of core ai platform
  • Assets built on it

Job Description:

  • Client is expanding it s Global business in 2026 and the Cloud and Infrastructure team will need to support this growth by designing, provisioning then supporting the platforms to enable this.
  • We are seeking an experienced Site Reliability Engineer (SRE) to join our Cloud & Infrastructure team. The successful candidate will be responsible for designing, operating, automating, and continuously improving enterprise-scale Azure platforms, ensuring high availability, resiliency, security, and performance.
  • The role combines software engineering, cloud architecture, infrastructure automation, AI, and operational excellence to improve service reliability and reduce operational overhead through automation and engineering best practices.
  • The ideal candidate will have strong experience with Microsoft Azure, DevOps tools & practices, Terraform, API, AI platforms & tools, disaster recovery planning, and enterprise-scale resilience engineering.

Key Responsibilities

Platform Reliability & Operations

  • Ensure the availability, performance, scalability, and reliability of Azure-hosted services.
  • Define and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Proactively monitor platform health and performance using observability tooling.
  • Perform root cause analysis and implement permanent fixes for recurring incidents.
  • Participate in incident management and on-call support rotations where required.
  • Lead blameless post-incident reviews, capture lessons learned, and drive corrective actions through to completion.
  • Reduce operational toil by identifying repetitive manual tasks and replacing them with automated, reusable engineering solutions.
  • Develop reliability dashboards and actionable alerts that focus on customer-impacting symptoms rather than infrastructure noise.

Azure Cloud Engineering

  • Design, deploy, and manage Azure infrastructure services including:
    • Virtual Networks
    • Application Gateways
    • API Management
    • Azure Kubernetes Service (AKS)
    • Azure Firewall
    • Azure Storage
    • Key Vault
    • Azure Monitor
    • Azure AI Services
  • Implement cloud platform standards and best practices.
  • Support multi-region Azure deployments and platform modernisation initiatives.
  • Undertake capacity planning and performance engineering to ensure platforms can scale reliably in line with business growth and peak demand.

Infrastructure as Code (Terraform)

  • Develop and maintain Terraform modules and reusable infrastructure patterns.
  • Implement Infrastructure as Code (IaC) standards and governance controls.
  • Ensure infrastructure is version controlled, peer-reviewed, and fully automated.
  • Manage Terraform state securely and consistently across environments.

DevOps, Automation & AI

  • Build and maintain Azure DevOps CI/CD pipelines.
  • Automate infrastructure provisioning and application deployments using pipelines with automated delivery and testing routines.
  • Implement testing, security scanning, policy compliance, and release gates.
  • Support DevOps and platform engineering practices.
  • Create automation for operational runbooks, self-healing processes, deployment validation, and environment consistency checks.
  • Manage and Implement AI platforms and tools such as Claude & Open AI, to develop skills and support business adoption of agentic AI capabilities.
  • Implement and enable self-service approach to technology services.

Resilience, Disaster Recovery & Failover

  • Design and implement highly available Azure architectures.
  • Develop and maintain disaster recovery and business continuity capabilities.
  • Implement and test:
    • Regional failover strategies
    • Active/Passive architectures
    • Active/Active deployments
    • Traffic Manager and Front Door failover patterns
    • Database resiliency and replication
    • Backup and recovery solutions
  • Conduct regular resilience and recovery testing exercises.
  • Identify and reduce single points of failure across platforms.
  • Define and execute game days, chaos testing, and controlled failure scenarios to validate operational resilience.

Security & Governance

  • Ensure platforms are secure-by-design.
  • Work closely with Security and Architecture teams to implement:
    • Zero Trust principles
    • RBAC controls / Managed Identities
    • Network segmentation and Zone based architecture
    • Secrets management
  • Support compliance requirements and operational audits.
  • Help coordinate security updates, patches, maintenance routines, and upgrades of the underlying system across partners and vendors
  • Embed reliability, security, and compliance controls into build and release pipelines to support production readiness.

Continuous Improvement

  • Drive automation and reduction of manual operational tasks.
  • Improve deployment reliability and platform observability.
  • Contribute to architecture standards, runbooks, and operational documentation.
  • Partner with engineering, architecture, security, and service teams to define production readiness standards and reliability acceptance criteria.

About You

At Client we work in a fast paces evolving environment with a growth mindset and outcome focus.

Core Skills & Experience

  • Microsoft Azure Expert
    • Azure Administration and Architecture
    • Networking and Connectivity
    • Virtual Machines and Platform Services
    • Azure Monitor and Log Analytics
    • Azure Identity and Access Management
    • Azure Networking, NSGs, Firewalls, Load Balancing
    • Azure Backup and Disaster Recovery
  • Strong expertise in Azure networking (VNets, routing, firewalls, private links, load balancing).
  • Hands-on proficiency with infrastructure-as-code and automated deployments. (Must have Terraform and Git Enterprise, orchestration engines)
  • Exposure and understanding of building, deploying and managing API Gateways
  • Strong understanding of Azure security controls, governance, and compliance frameworks.
  • Full stack observability e.g. MELTS principles golden signals, and automation response using DataDog, New Relic, Splunk or other leading tools.
  • Strong FinOps expertise
  • Scripting skills (PowerShell, Bash, Python, React).
  • Strong understanding of Devops practices, tooling, and SDLC methods
  • Strong exposure to Anthropic, Open AI, platforms and associated tools & practices e.g Harness, Token usage, Skills, LLM and SLM concepts, Orchestration engines, and agent cost management.
  • Strong understanding of Site Reliability Engineering principles, including SLIs, SLOs, SLAs, error budgets, reliability targets, and service health measurement.
  • Experience designing observability strategies across metrics, logs, traces, synthetic monitoring, alerting, dashboards, and operational telemetry.
  • Ability to define actionable alerts that identify customer-impacting symptoms, reduce noise, and support rapid incident triage.
  • Proven capability in incident response, root cause analysis, blameless post-incident reviews, corrective action tracking, and operational learning.
  • Experience reducing toil through automation, self-service tooling, runbook automation, self-healing patterns, and repeatable engineering solutions.
  • Strong knowledge of capacity planning, performance engineering, load testing, scalability modelling, saturation analysis, and demand forecasting.
  • Experience with resilience validation techniques including chaos engineering, game days, failover testing, disaster recovery exercises, and operational readiness testing.
  • Ability to establish production readiness standards, reliability acceptance criteria, operational runbooks, service ownership models, and support handover practices.
  • Working knowledge of deployment reliability practices such as canary releases, blue-green deployments, rollback strategies, feature flags, and release health monitoring.

Thanks & Regards

Saravanan

DMinds Solutions Inc.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10513378
  • Position Id: 9074629
  • Posted 55 minutes ago
Contact the job poster
SM

saravanan Muthusamy

Recruiter @ Dminds Solutions Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Easy Apply

Third Party, Contract

Depends on Experience

Remote

Today

Full-time

USD 115,000.00 - 160,000.00 per year

Remote

21d ago

Easy Apply

Contract

Depends on Experience

Remote

Today

Full-time

USD 151,500.00 - 252,500.00 per year

Search all similar jobs