Lead Site Reliability Engineer (F2F - NY)

Buffalo, NY, US • Posted 14 hours ago • Updated 14 hours ago
Contract Corp To Corp
Contract W2
Contract Independent
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

⏳ Almost there, hang tight...

Job Details

Skills

  • SRE
  • DYNATRACE
  • OPENTELEMETRY
  • Infrastructure as Code
  • Microsoft Azure
  • Azure App Services
  • Service Level Objectives (SLOs)
  • Reliability automation

Summary

Job Title : Lead Site Reliability Engineer

Location: Buffalo, NY 100% onsite

Duration: Long Term

Primary Responsibilities

  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.
  • Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.
  • Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.
  • Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.
  • Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
  • Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.
  • Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.
  • Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
  • Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).
  • Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
  • Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
  • Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.
  • Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.
  • Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.
  • Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.
  • Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.
  • Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.
  • Lead capacity planning, performance tuning, and workload optimization efforts across production environments.
  • Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
  • Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.
  • Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
  • Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.
  • Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.
  • Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite.
  • Identify reliability, operational, and technology risks requiring escalation to management.
  • Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
  • Maintain M&T internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
  • Complete other related duties as assigned.

Core Requirements

  • Strong experience in observability and monitoring, including hands-on expertise with:
  • Dynatrace
  • OpenTelemetry (OTel)
  • Distributed tracing
  • Metrics collection and analysis
  • Centralized logging and log aggregation
  • Alerting and dashboard development
  • Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform.
  • Experience with CI/CD pipelines, deployment automation, and operational tooling.
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
  • Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.

Cloud & Platform Expertise

  • Strong experience with Microsoft Azure, including:
  • Azure App Services
  • Resource Groups
  • Azure networking concepts
  • Scaling and performance optimization
  • Deployment and release management
  • Application lifecycle management
  • Experience leveraging Azure-native operational tooling such as:
  • Azure Monitor
  • Application Insights
  • Log Analytics
  • Azure dashboards and alerting
  • Experience supporting cloud-native and hybrid infrastructure environments.

Reliability & Engineering Practices

  • Demonstrated experience implementing and operating SRE practices, including:
  • Service Level Objectives (SLOs)
  • Service Level Indicators (SLIs)
  • Error budgets
  • Incident management
  • Problem management
  • Root Cause Analysis (RCA)
  • Reliability automation
  • Ability to improve system reliability through:
  • Performance tuning
  • Capacity planning
  • Observability-driven insights
  • Proactive issue detection
  • Reliability engineering initiatives
  • Experience developing automated recovery mechanisms and self-healing solutions.
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91165791
  • Position Id: 9049456
  • Posted 14 hours ago

Company Info

About Keen Technology Solutions LLC

Keen Technology Solutions headquarted in Irving, TX, with more than 25 Years of IT professional resourcing and managed services, we empower organizations across the full lifecycle of technology adoption, providing flexible solutions that adapt to their needs as technologies and skills evolve.

Keen is a world-class technology services business that incorporates industry insights and experience to deliver solutions that fulfil our clients’ digital visions. We use our unique deployment model to build qualified, industry specialized fit-for-purpose teams combined with proven solutions and service models to achieve results. Our agility and obsession with providing value enables us to support an ever-evolving digital world.

WORKING WITH KEEN
"One big positive for me is Keen ensures there is a continuous dialogue with the client and wants to make sure the candidates are qualified and progressing as expected…. they are professionals with a strong customer focus. Keen Account Managers go the extra mile to ensure qualified candidates are presented and are always keeping in touch. That is important in this business. "

About_Company_OneAbout_Company_Two
Contact the job poster
VV

Vinay Varma

Recruiter @ Keen Technology Solutions LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs