Senior Site Reliability Engineer

Overview

On Site
$60-65
Accepts corp to corp applications
Contract - 20 day((s))
50% Travel

Skills

AWS
python
Prometheus

Job Details

Senior Site Reliability Engineer
Position: Senior Site Reliability Engineer
Location: Owings Mills, MD (No relocation candidates)
2 days onsite/3 days remote
Interview: 1st round virtual, 2nd round onsite
Role summary and job responsibilities
Design and implement highly automated systems/services that ensure the availability, reliability, and scalability of infrastructure and applications.
Build and maintain monitoring and alerting to provide timely feedback on the performance and health of systems, network, and applications.
Design and implement automation tools to reduce manual toil, streamline repetitive tasks, and enhance overall operational efficiency.
Design and build Service Level Indicator (SLIs) metrics, including but not limited to Service Level Objectives (SLOs), Error Budget, Burn Rate Alerts
Work closely with development teams to embed reliability best practices into the software development process. Provide mentorship and training to cross-functional teams on SRE principles, encouraging a shared responsibility for the reliability of our services.
Collaborating with our support, operations and engineering teams to investigate and troubleshoot complex problems
Observe and monitor systems to make sure you have the insight into system performance, health, availability and what is happening internally in the system.
Understands what to monitor based on the system(s) you are managing, how the monitoring data is stored, and how to look at the data to make determinations about future actions.
Participates in continuous improvement efforts that span multiple multi-functional domains and informs the generation of new standards
Be a part of an on-call rotation, continuously enhance automation & documentation, and mentor others on the standard methodologies of infrastructure automation to encourage adoption.
Able to overcome differences of opinion and drive team alignment around a specific goal or solution
Holds associates and teams accountable for adhering to practices and policies
Requirements
Strong experience with Monitoring and Alerting tools such as Prometheus, Grafana, New Relic
Experience in container orchestration solutions in AWS with ECS, Fargate
Docker container development experience
Scripting languages like Python, Groovy, PowerShell, Bash, Perl etc.
Skilled in building and maintaining dashboards using tools like Grafana, Prometheus and Statsd to provide critical insights
Worked with Service Reliability Engineering team to design SLI and SLO for respective applications
Strong experience with AWS cloud infrastructure and container orchestration operating in a GitOps framework
A solid core foundation in infrastructure and systems engineering including Unix/Linux compute, networking, storage, and monitoring stacks.
Have experience using automation tools such as Terraform, Ansible
Excellent written and oral communication skills
Strong interpersonal skills, adaptable and able to learn quickly
Off-hour implementations are required
Ability to build positive working relationships with the business contacts, within our IT team, and other IT departments
Ability to identify tasks and help develop project plans for medium and large-scale projects
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.

About AspireIT Solutions