Systems Reliability Engineer

Berkeley Heights, NJ, US • Posted 3 days ago • Updated 3 days ago
Contract W2
12 Weeks
Travel Required
Able to Sponsor
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Programming Languages
  • Embedded Systems
  • MEAN Stack
  • Finance
  • Financial Services
  • Continuous Delivery
  • Incident Management
  • IT Operations
  • Management
  • Budget
  • Build Automation
  • Dynatrace
  • Continuous Integration
  • Bash
  • Collaboration
  • Communication
  • Reliability Engineering
  • Grafana
  • Scripting
  • Root Cause Analysis
  • Service Level
  • Cloud Computing
  • Payments
  • Python
  • ROOT
  • Operational Efficiency
  • ServiceNow
  • Splunk
  • Reporting
  • Workflow

Summary

Experience: 8 To 12 Years
Job Title: Systems Reliability Engineer, Embedded Finance
 

Location: Jacksonville, FL; Berkeley Heights, NJ; Alpharetta, GA, Toronto, ON  (on-site)

Job Description: 

What does a successful Systems Reliability Engineer (SRE) in Embedded Finance ?

We are seeking a Systems Reliability Engineer to join the Technical Operations team supporting our Embedded Finance (EmFi) platform. In this role, you will own the reliability and resiliency of a large-scale enterprise platform, ensuring that our services remain highly available, performant and secure. You will design and implement monitoring and alerting frameworks, lead incident response and drive the root cause analysis (RCA) process to continuously improve platform stability. This is an opportunity to be a Partner in Possibility — helping our clients deliver financial services experiences that are essential to everyday life. 

What you will do: 

  • Own the reliability, resiliency and availability of the Embedded Finance platform, proactively identifying and mitigating risks to service continuity. 
  • Design, implement and maintain comprehensive monitoring and alerting frameworks leveraging Splunk, Dynatrace, Grafana and Datadog to provide end-to-end observability across the platform. 
  • Define and track service level objectives (SLOs), service level indicators (SLIs) and error budgets to measure and improve platform health. 
  • Lead and participate in incident response, serving as a technical driver during remediation calls and coordinating with impacted and impacting technical and product teams. 
  • Own and advance the root cause analysis (RCA) process — investigating incidents, documenting the sequence of events and remediating actions, and clearly identifying underlying root causes to prevent recurrence. 
  • Ensure timely creation and management of incident tickets (e.g., ServiceNow) and accurate incident tracking, aging and reporting. 
  • Build automation and tooling to reduce toil, improve mean time to detection (MTTD) and mean time to resolution (MTTR), and increase operational efficiency. 
  • Collaborate with engineering, product and risk stakeholders to embed reliability best practices into the platform lifecycle. 

What you will need to have: 

  • Hands-on experience with monitoring, observability and alerting tools, specifically Splunk, Dynatrace, Grafana and Datadog. 
  • Proven experience operating and supporting a large-scale enterprise platform environment. 
  • Demonstrated experience with incident response and leading or contributing to root cause analysis (RCA) processes. 
  • Strong understanding of reliability engineering principles, including availability, resiliency, monitoring and alerting best practices. 
  • Experience with ticketing and incident management workflows (e.g., ServiceNow). 
  • Excellent communication skills, with the ability to drive remediation efforts and collaborate across technical, product and risk teams. 

What would be great to have: 

  • Experience in financial services, payments or embedded finance environments. 
  • Proficiency with scripting or programming languages for automation (e.g., Python, Go, Bash). 
  • Familiarity with cloud platforms, containerization and CI/CD pipelines. 
  • Experience defining and managing SLOs, SLIs and error budgets.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10495938
  • Position Id: 9092419
  • Posted 3 days ago

Company Info

About Ekcel Technologies Inc

Ekcel was started with an idea of recreating the services and solutions in the area of Telecom, IT and Networking with a new prospect. We at Ekcel believe innovation is the way moving forward to cater the ever increasing demand from the shrinking global market.

At Ekcel, we endeavor to achieve and perfect exactly this- implementing solutions for a changing world. A glimpse of the future some may say. Achieving this is not an easy task. It requires a highly dedicated and talented engineering workforce with deep knowledge, multiple skill-set and the right attitude.

Contact the job poster
Rajesh Khanna

Rajesh Khanna

Recruiter @ Ekcel Technologies Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs