Reliability engineer

Hybrid in Dallas, TX, US • Posted 12 hours ago • Updated 12 hours ago
Full Time
No Travel Required
Hybrid
$65 - $68/yr
Fitment

Dice Job Match Score™

⭐ Evaluating experience...

Job Details

Skills

  • Kubernetes
  • Python
  • PostgreSQL
  • golang
  • aws
  • azure
  • Grafanais
  • Amazon Web Services
  • Automation

Summary

The Role

As Senior Reliability Engineer, you blend deep Infrastructure automation experience and reliability engineering expertise with a passion for delivering results. Our Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high availability with resilience by using automation and Infrastructure Code. We build reliability into our ecosystem by applying standards in Resiliency Engineering, Automation, Observability & Chaos Testing.

Additionally, this role contributes to enterprise backup and recovery capabilities including automation of recovery workflows, rehoused recovery into alternate datacenters, and testing of recovery processes.

We are looking for a systems thinking, reliability engineer who has helped teams scale through production insight, data and backup recovery, operational automation, developer guidance, real-time metrics, automation, automation, automation.

  • Strong background in several of the following: Go, Angular, Python, JavaScript, AWS, RESTful services, Ruby, MVC, Jenkins CI/CD, Configuration Automation (Chef, Ansible).
  • Preferred background in: Bootstrap, HTML/CSS, Shell Scripting, messaging frameworks (MQ), Service Oriented/Micro-service Architectures, OpenStack, Relational Databases (PostgreSQL).
  • Comfortable working in both Public and private cloud environments.
  • Crafting scalable solutions and automation to monitor the health and establish signals to drive understanding of our Container Platform environments.
  • Strengthening operational processes with support and incident management teams for our cloud ecosystem
  • Working with and cloud service provider product teams and driving ongoing reliability improvements in their Kubernetes service offerings.
  • Anticipating, discovering through ongoing interaction with, and prioritizing client / partner needs to serve as their voice and guide execution of the team.

The Expertise You Have

  • Bachelor’s Degree or equivalent experience in a technology related field (e.g. Computer Science, Engineering, etc.) required.
  • Production experience running Cloud and on-prem Storage workloads at scale
  • Experience managing and maintaining Kubernetes Clusters on EKS/AKS and RKS.
  • Demonstrates a drive for continuous improvement and enjoys tackling complex problems.
  • Experience managing and interpreting large datasets using query languages and visualization tools(PowerBI/tableau),
  • Experience in software development with Python, NodeJS, or Java with a focus on SDLC and automation
  • 5 -7 years of hands-on experience deploying and/or supporting highly distributed multi-tiered systems at scale.
  • Experience building and deploying Docker images including Docker Compose
  • Hands-on experience with Jenkins Core, including authoring and maintaining declarative CI/CD pipelines and libraries
  • Experience with distributed version control systems, Git preferred
  • Experience crafting and maintaining logging, monitoring, and alerting capabilities using tools like Datadog and Splunk
  • Practical experience in building cloud hosted and native applications for the enterprise. Maintains a deep understanding of a wide variety of AWS/Azure services that support reliability, observability, and automation/orchestration.
  • Experience in incident/crisis management and supporting critically important applications

The Skills You Bring

  • Hands on experience with one or more observability tools (Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, Datadog, etc.)
  • Ability to automate with various scripting languages (Python, Shell scripting, etc.)
  • Experience managing systems using infrastructure as code tools (IAM, ARM, Terraform, Chef)

Additional Value in Backup & Recovery:

  • Advance enterprise resiliency through improved recovery capabilities.
  • Reduce recovery time via automation.
  • Enable rehoused recovery into new datacenters.
  • Strengthen platform reliability through data protection design.

Please be advised that ’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict from hiring and/or associating with individuals with certain Criminal Histories.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10227485
  • Position Id: 9046478
  • Posted 12 hours ago
Contact the job poster
DS

Divya Subramani

Recruiter @ Digipulse Technologies, Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Dallas, Texas

16d ago

Easy Apply

Full-time

$170,000 - $190,000

Texas

Today

Contract

USD67 - USD68

Remote

Today

Full-time

USD 205,000.00 - 270,000.00 per year

No location provided

Today

Full-time

USD 230,000.00 - 250,000.00 per year

Search all similar jobs