Engineer, Site Reliability

• Posted 3 days ago • Updated 2 days ago
Full Time
On-site
Fitment

Dice Job Match Score™

🛠️ Calibrating flux capacitors...

Job Details

Skills

  • IT Strategy
  • Reliability Engineering
  • Customer Facing
  • Stacks Blockchain
  • Failover
  • Recovery
  • Training
  • Systems Architecture
  • Scalability
  • Performance Tuning
  • IT Management
  • Java
  • JavaScript
  • Cloud Computing
  • Microservices
  • Root Cause Analysis
  • Amazon Web Services
  • Cloud Architecture
  • Python
  • Scripting
  • Workflow
  • Software Engineering
  • Finance
  • Collaboration

Summary

Core Responsibilities


  • Lead the technical strategy, architecture, and evolution of PWTech reliability engineering platforms and capabilities, ensuring they scale across hundreds of applications and critical client-facing systems.

  • Design and build production-grade software and platforms that improve reliability outcomes, including automated incident detection, diagnostics, remediation, resiliency engineering, and operational intelligence.

  • Drive enterprise observability and diagnostics capabilities, enabling consistent telemetry, distributed tracing, metrics, and operational insights across cloud-native technologies and application stacks.

  • Define and codify resilient application and platform patterns, such as graceful degradation, circuit breakers, load shedding, fault isolation, failover, and automated recovery, driving adoption through reusable software, frameworks, and engineering standards.

  • Influence engineering teams and technology leaders across the organization, establishing technical standards and ensuring reliability is designed into systems from inception.

  • Lead the resolution of complex reliability and production challenges, identifying systemic risks, driving root cause analysis, and engineering durable solutions that improve long-term resilience.

  • Participates in special projects and performs other duties as assigned.



Qualifications


  • Minimum of eight years related experience, with at least two years of development experience.

  • Undergraduate degree or equivalent combination of training and experience. Graduate degree preferred.



Preferred Skills


  • Experience designing, building, and operating production-facing platforms or engineering capabilities that achieve broad adoption and deliver measurable reliability, operational, or business outcomes.

  • Deep expertise in distributed systems architecture, including scalability, availability, resiliency, fault tolerance, performance optimization, and production operations at scale.

  • Strong technical leadership and influence skills, with a demonstrated ability to drive architecture decisions, establish technical standards, and align multiple teams on engineering direction.

  • Deep expertise in Java or JavaScript, with hands-on experience developing and operating software in modern cloud-native and microservices environments.

  • Demonstrated ability to diagnose and resolve complex production issues, perform root cause analysis, and engineer durable solutions that prevent recurrence.

  • Hands-on experience with AWS and modern cloud architecture patterns.

  • Experience with observability and telemetry platforms, including metrics, logging, distributed tracing, and production diagnostics. Experience with OpenTelemetry is strongly preferred.

  • Proficiency with Python or similar scripting languages to develop automation, tooling, and operational workflows.

  • Strong software engineering fundamentals, systems thinking skills, experience solving complex reliability challenges, and the ability to influence across teams and drive engineering best practices



Special Factors

Sponsorship

Vanguard is not offering visa sponsorship for this position.

About Vanguard

At Vanguard, we don't just have a mission-we're on a mission.

To work for the long-term financial wellbeing of our clients. To lead through product and services that transform our clients' lives. To learn and develop our skills as individuals and as a team. From Malvern to Melbourne, our mission drives us forward and inspires us to be our best.

How We Work

Vanguard has implemented a hybrid working model for the majority of our crew members, designed to capture the benefits of enhanced flexibility while enabling in-person learning, collaboration, and connection. We believe our mission-driven and highly collaborative culture is a critical enabler to support long-term client outcomes and enrich the employee experience.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90922487
  • Position Id: 24660615
  • Posted 3 days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Malvern, Arkansas

2d ago

Full-time

Montana

Today

Full-time

USD 189,000.00 - 232,000.00 per year

Remote

Today

Full-time

USD 147,000.00 - 168,000.00 per year

Kentucky

Today

Full-time

Search all similar jobs