Lead Site Reliability Engineer - Infrastructure & DevOps

Orlando, FL, US • Posted 11 hours ago • Updated 11 hours ago
Full Time
No Travel Required
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🧠 Analyzing your skills...

Job Details

Skills

  • Kubernetes
  • Terraform
  • Helm
  • Aws
  • Azure
  • SRE
  • DevOps

Summary

Job Title: Lead Site Reliability Engineer (SRE)

Overview / Summary

We are seeking a Lead Site Reliability Engineer to help drive the reliability, scalability, and operational excellence of a rapidly growing Generative AI platform. This role provides technical leadership while designing and supporting highly available cloud infrastructure powering modern AI and data-driven applications.

The ideal candidate combines deep expertise in Site Reliability Engineering, cloud infrastructure, Kubernetes, and Infrastructure as Code with strong leadership skills. You will work alongside platform engineers, architects, and development teams to build resilient systems, improve automation, and ensure high availability across a multi-cloud environment.

Key Responsibilities

  • Lead the design, implementation, and support of highly available cloud infrastructure across Google Cloud Platform (primary), AWS, and Azure.
  • Design, build, and maintain Kubernetes infrastructure using Helm and Terraform for Infrastructure as Code.
  • Develop scalable platform solutions capable of maintaining 99.99% service availability.
  • Lead and mentor Site Reliability Engineers and DevOps engineers by providing technical guidance and establishing engineering best practices.
  • Plan, prioritize, and coordinate infrastructure initiatives within Agile delivery teams.
  • Design and implement automated deployment pipelines using modern CI/CD tools, including Harness.
  • Implement progressive deployment strategies such as blue/green deployments, canary releases, and feature flag rollouts.
  • Build and enhance observability solutions using monitoring, logging, alerting, and distributed tracing technologies.
  • Partner with engineering teams to review infrastructure sizing, capacity planning, and scalability requirements.
  • Support production systems through backups, upgrades, patching, disaster recovery, and operational maintenance.
  • Troubleshoot complex production issues across distributed systems and cloud-native applications.
  • Evaluate emerging SRE and DevOps technologies and recommend improvements to platform reliability and operational efficiency.
  • Ensure infrastructure aligns with security, governance, and compliance standards.

Required Qualifications

  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles.
  • Expert-level experience administering and operating Kubernetes in production environments.
  • Strong experience with Helm for Kubernetes application management.
  • Advanced experience using Terraform for Infrastructure as Code.
  • Hands-on experience building automated deployment pipelines using Harness or comparable enterprise CI/CD platforms.
  • Experience supporting production workloads across Google Cloud Platform, AWS, and Azure.
  • Strong scripting and automation skills using Python, Bash, and YAML.
  • Experience supporting production databases and messaging technologies, including PostgreSQL, Redis, Kafka, MongoDB, and Vault.
  • Experience with enterprise CI/CD platforms such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or Harness.
  • Experience implementing observability solutions using technologies such as OpenTelemetry, Prometheus, Splunk, AppDynamics, or similar platforms.
  • Strong troubleshooting skills within distributed systems and cloud-native environments.
  • Experience working within Agile development environments.

Excellent communication skills with the ability to explain complex technical concepts to both technical and non-technical audiences

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10354711
  • Position Id: 9052635
  • Posted 11 hours ago
Contact the job poster
Praveen Palavalasa

Praveen Palavalasa

Lead Talent Acquisition @ SRI Tech Solutions
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Lake Buena Vista, Florida

Today

Easy Apply

Contract

$83 - $90

Lake Mary, Florida

Today

Full-time

Remote

4d ago

Easy Apply

Full-time, Third Party

Depends on Experience

Remote

Today

Easy Apply

Full-time

$80000 - $120000

Search all similar jobs