Product Reliability Engineering Lead

JERSEY CITY, NC, US • Posted 9 hours ago • Updated 36 minutes ago
Full Time
Part Time
On-site
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • RELIABILITY ENGINEERING
  • CI/CD

Summary

Client: Corebridge

Job Title: Product Reliability Engineering Lead

Location: Houston, TX / Jersey City, NJ / Raleigh, NC / Remote

Visa: L2S

Experience: 5+ Years

Employment Type: Contract

C2C Rates

Houston, TX: $65/hr C2C

Jersey City, NJ: $70/hr C2C

Raleigh, NC: $60hr C2C

Remote: $60/hr C2C

About the Role



Corebridge is transforming its digital product offering across the insurance and retirement value chain. The Acquisition team is building a next-generation platform for the submission and processing of new applications for insurance and retirement products.

The Product Reliability Engineering Lead will be responsible for ensuring that the platform scales reliably and sustainably as demand grows across products, channels, and distribution partners.

This role focuses on shifting reliability left into design and development, establishing observability as a core platform capability, and using AI-assisted techniques to accelerate detection, diagnosis, and resolution of issues before they affect partners or customers.

Required Experience

5+ years of experience in Site Reliability Engineering, Platform Engineering, or Production Operations.

Experience defining and operating SLO/SLI frameworks tied to business outcomes.

Hands-on experience designing observability for distributed, API-driven platforms.

Experience with reliability and resiliency testing, including chaos engineering and fault injection.

Experience guiding and mentoring engineers on reliability practices.

Enterprise-scale delivery experience working with both onshore and offshore cross-functional teams.

Direct experience with Agile methodologies in product-centric delivery models.

Preferred Experience

Financial services, insurance, or retirement services industry experience.

AWS operational experience with:

CloudWatch

X-Ray

Fault Injection Simulator

ECS/EKS

Lambda

EventBridge

Experience integrating reliability practices with DevSecOps and CI/CD pipelines.

Familiarity with AI/ML-driven operations tools and incident management platforms.

Reliability Strategy & Architecture

Define and lead reliability strategies aligned with product, platform, and enterprise goals.

Establish SLOs, SLIs, and error budgets tied to business outcomes.

Shift reliability requirements into early design and development.

Design resiliency, failover, and graceful-degradation patterns.

Design reliability patterns across platform services, APIs, workflows, and internal/external dependencies.

Observability & Operational Enablement

Architect end-to-end observability covering:

Metrics

Structured logging

Distributed tracing

Alerting

APM telemetry

Establish monitoring standards and dashboards for platform health and partner-facing services.

Embed observability into services from design through deployment.

Drive adoption of synthetic monitoring and canary deployments.

AI-Assisted Reliability Engineering

Guide teams in using prompt-driven and agent-based approaches to automate operational toil.

Introduce AI-enabled patterns for predictive alerting, automated remediation, and intelligent escalation.

Apply responsible AI practices with consideration for security, data exposure, and operational risk.

Reliability Testing & Validation

Partner with development and test engineering teams to align functional, non-functional, and reliability testing.

Embed reliability validation into CI/CD pipelines.

Validate platform behavior under failure conditions across APIs, integrations, workflows, and experience layers.

Perform or support chaos engineering, fault injection, load/stress testing, and failure-mode analysis.

Technical Skills

SRE principles

SLOs, SLIs, error budgets

Toil reduction and blameless postmortems

Observability and APM

Distributed tracing

Structured logging

Real-time alerting

Synthetic monitoring

Chaos engineering

Fault injection

Resilience and fault tolerance

Circuit breakers

Bulkheads

Retry/backoff

Graceful degradation

Failover validation

CI/CD reliability integration

Automated reliability gates

Canary deployments

Feature flags

Progressive rollouts

AI-assisted reliability techniques

Anomaly detection and predictive alerting

Prompt-driven runbook automation

Agent-based remediation

Cloud-native operations

Containerized platforms

Event-driven architectures

Infrastructure as Code

Soft Skills

Strong communication and stakeholder-management skills.

Ability to translate reliability concerns between engineering, product, and business teams.

Strong leadership, mentoring, and collaboration skills.

Growth-oriented mindset with the ability to identify future reliability needs.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91171926
  • Position Id: OOJ - 1814-818-1790100322
  • Posted 9 hours ago

Company Info

About StratEdge It consulting INC

We are a specialized IT consulting firm dedicated to providing strategic solutions and robust technical support tailored to meet your organization's unique technology needs. With deep expertise across cloud infrastructure, networking, cybersecurity, and systems management, our team helps businesses optimize their technology landscape, enhance operational efficiency, and drive innovation.

Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs