Product Reliability Engineering Lead

Remote • Posted 11 hours ago • Updated 11 hours ago
Contract W2
Contract Independent
24 Months
No Travel Required
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • API
  • Acquisition
  • Amazon Web Services
  • Agile
  • Communication
  • Budget
  • Artificial Intelligence
  • Cloud Computing
  • CHAOS
  • Continuous Delivery
  • Collaboration
  • Continuous Integration
  • DevSecOps
  • Dashboard
  • Failover
  • Distribution
  • Financial Services
  • Incident Management
  • MEAN Stack
  • Jersey
  • Insurance
  • Machine Learning (ML)
  • Mentorship
  • Management
  • Operational Risk
  • Offshoring
  • Real-time
  • Pivotal
  • Reliability Engineering
  • Recovery
  • Software Performance Management
  • Stakeholder Engagement
  • Testing
  • Stress Testing
  • Workflow

Summary

 
 

Description: Product Reliability Engineering Lead 

 

Locations: Houston, Jersey City, Raleigh, Remote

 

Max Rates:

About the role
What you need to know:

Client is transforming our digital product offering across the insurance and retirement value chain. Our Acquisition team is delivering a next generation platform for the submission and processing of new applications for our insurance and retirement products. This platform will help automate and reduce manual work to shorten the time to issue a new policy/contract to customers.
The Platform Reliability Engineering Lead for the Acquisition Platform is a pivotal role responsible for ensuring the platform scales reliably and sustainably as demand grows across products, channels, and distribution partners. This role shifts reliability left into design and development, embeds observability as a first‑class platform capability, and leverages AI‑assisted techniques to accelerate detection, diagnosis, and resolution of issues before they impact partners or customers.
The ideal candidate brings deep experience in site reliability engineering, platform observability, and resilience validation, along with the ability to lead teams in treating reliability as a continuous, measurable product discipline rather than a reactive operations function.

What we’re looking for:

 

Required Experience
• 5+ years of experience in site reliability engineering, platform engineering, or production operations roles
• Experience defining and operating SLO/SLI frameworks tied to business outcomes
• Hands‑on experience designing observability for distributed, API‑driven platforms
• Experience with reliability and resiliency testing including chaos engineering and fault injection
• Experience guiding and mentoring engineers on reliability practices
• Enterprise‑scale delivery experience with both onshore and offshore cross‑functional teams
• Direct experience applying Agile methodologies in product‑centric delivery models

Preferred
• Experience in financial services, insurance, or retirement services industry
• AWS operational experience – CloudWatch, X‑Ray, Fault Injection Simulator, ECS/EKS, Lambda, EventBridge
• Experience integrating reliability practices with DevSecOps and CI/CD pipelines
• Familiarity with AI/ML‑driven operations tools and incident management platforms


Reliability Strategy and Architecture
• Define and lead the reliability strategy for the Acquisition Platform, ensuring alignment with product, platform, and enterprise goals.
• Establish SLOs, SLIs, and error budgets that tie reliability targets to business outcomes and partner expectations.
• Shift reliability requirements into early design and development phases so resiliency, failover, and graceful degradation are architected in, not bolted on.
• Design reliability patterns across platform services, APIs, workflows, and dependent systems both internal and external to Corebridge.

Observability and Operational Enablement
• Architect end‑to‑end observability across the platform including metrics, structured logging, distributed tracing, and alerting.
• Establish monitoring standards and dashboards that provide real‑time visibility into platform health, partner‑facing services, and integration dependencies.
• Embed observability into platform services from design through deployment so teams can detect, diagnose, and resolve issues rapidly.
• Drive adoption of synthetic monitoring and canary deployments to validate production behavior proactively.
AI‑Assisted Reliability Engineering
• Direct teams in using prompt‑driven and agent‑based approaches to automate toil, reduce mean time to recovery, and improve operational consistency.
• Explore and introduce AI‑enabled patterns for predictive alerting, automated remediation, and intelligent escalation where appropriate.
• Apply responsible AI practices with attention to security, data exposure, and operational risk.

Reliability Testing and Validation
• Partner with the Platform Development Engineer in Test to align functional, non‑functional, and reliability test coverage.
• Embed reliability validation into CI/CD pipelines so resiliency is continuously tested, not assumed.
• Validate platform behavior under failure conditions across APIs, integrations, workflows, and experience layers.
Collaboration and Stakeholder Engagement
• Collaborate closely with the Acquisition delivery team and stakeholders to align outcomes with the reliability strategy.
• Partner with AMS, infrastructure, and other tech teams to ensure clear ownership boundaries and smooth operational handoffs.

Technical Skills and Competencies
• SRE principles – SLOs, SLIs, error budgets, toil reduction, blameless postmortems
• Observability design – distributed tracing, APM telemetry, structured logging, real‑time alerting, synthetic monitoring
• Resilience and fault tolerance – circuit breakers, bulkheads, retry/backoff, graceful degradation, failover validation
• Chaos engineering and reliability testing – fault injection, load/stress testing, failure‑mode analysis
• CI/CD reliability integration – automated reliability gates, canary deployments, feature flags, progressive rollouts
• AI‑assisted reliability techniques – anomaly detection, predictive alerting, prompt‑driven runbook automation, agent‑based remediation
• Responsible AI use – including consideration of security, data exposure, and operational risk
• Cloud‑native operations – containerized platforms, event‑driven architectures, infrastructure as code
• Growth‑oriented mindset – ability to think beyond constraints of today and identify what is required to build the future
• Excellent communication skills – ability to translate reliability concerns between engineering, product, and business teams

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: PTP4pnSUJdNrxnq
  • Position Id: 9091423
  • Posted 11 hours ago
Contact the job poster
Srinath Reddy

Srinath Reddy

Recruiter @ METANLYTICS LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

15d ago

Full-time

USD 129,300.00 - 177,800.00 per year

Remote

Today

Full-time

USD 175,000.00 - 195,000.00 per year

Remote or Atlanta, Georgia

Today

Full-time

USD 120,000.00 - 175,000.00 per year

Remote

Today

Full-time

Search all similar jobs