Senior SRE / Cloud / Kubernetes / Terraform / 100% Remote

Remote in Remote, NJ, US • Posted 3 hours ago • Updated 3 hours ago
Full Time
On-site
$150000 - $200000/yr
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Business-to-business
  • Scalability
  • Product Engineering
  • Cloud Computing
  • Amazon Web Services
  • Google Cloud Platform
  • Google Cloud
  • Linux Administration
  • Computer Networking
  • Scripting
  • Python
  • Bash
  • Grafana
  • GPU
  • Machine Learning (ML)
  • Terraform
  • Continuous Integration
  • Continuous Delivery
  • Linux
  • Kubernetes
  • Incident Management
  • Reliability Engineering
  • IaaS
  • Management
  • Collaboration
  • Training And Development
  • Budget
  • Artificial Intelligence
  • SAP BASIS

Summary

Remote (USA) | Full-Time | Site Reliability Engineer
Join a rapidly growing B2B AI infrastructure company powering large-scale machine learning and AI workloads for more than one million developers worldwide. As a Site Reliability Engineer, you'll help improve the reliability, scalability, and performance of a cloud platform built on Linux, Kubernetes, distributed systems, GPU infrastructure, observability, and automation technologies. This full-time remote opportunity offers the chance to work on critical infrastructure supporting AI applications on a global scale.

As the company continues to scale its AI infrastructure platform, reliability has become a critical business function. This role sits at the center of that effort, partnering with Infrastructure, Product Engineering, and Support teams to improve uptime, strengthen observability, establish SLOs, reduce operational toil through automation, and lead incident response initiatives. The ideal candidate brings experience supporting large-scale production environments and enjoys solving complex reliability challenges while influencing engineering practices across a rapidly growing organization. This is an opportunity to gain exposure to cutting-edge AI and GPU infrastructure, take ownership of high-impact initiatives, and help shape the reliability strategy of a platform relied upon by more than one million developers.

Required Skills & Experience
5+ years of experience within major public cloud environment like AWS, Google Cloud Platform
Strong Linux systems administration experience
Strong networking fundamentals and troubleshooting skills
Experience supporting containerized environments (Kubernetes preferred)
Experience with monitoring, alerting, and observability tools
Experience defining and managing SLIs, SLOs, and reliability metrics
Incident response and postmortem experience
Scripting or programming experience. Python, Go, Bash, or similar technologies
Distributed systems and failure scenarios

Desired Skills & Experience
Kubernetes
Prometheus, Grafana, or similar monitoring platforms
Experience supporting GPU infrastructure or AI/ML platforms
Infrastructure as Code experience (Terraform preferred)
CI/CD pipeline experience

What You Will Be Doing
Tech Breakdown
40% Linux & Kubernetes Administration
25% Monitoring, Observability & Incident Response
20% Automation & Reliability Engineering
15% Distributed Systems & Cloud Infrastructure

Daily Responsibilities
80% Hands-On Engineering
5% Management Duties
15% Team Collaboration

The Offer
medical, dental, and vision benefits
Equity / Stock Options
Remote equipment stipend
Annual learning and development budget
Flexible PTO
Career Growth Within a Rapidly Scaling AI Infrastructure Company

Applicants must be currently authorized to work in the US on a full-time basis now and in the future. Sponsorship is not available for this position
#LI-JG2
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10105282
  • Position Id: 887531
  • Posted 3 hours ago

Company Info

About Motion Recruitment Partners, LLC

Motion Recruitment delivers IT Talent Solutions for Contract, Direct Hire, Managed Solutions and Statement of Work to all of North America from our 21 delivery centers. Our high-touch, specialized, team-based recruitment model’s success is proven through our exemplary track record in filling the most challenging IT positions for startup and enterprise clients alike. Our hyper-specialized tech focus results in a truly consultative approach for both our clients and candidates, within our recruiting areas of expertise: Software, Mobile, Data, Infrastructure, Cybersecurity, Product + UX and Functional.

Motion also delivers IT Consulting Solutions through the Motion Consulting Group (MCG) that create true digital transformation for IT projects in Agile Development & Coaching, DevOps & DevSecOps Solutions, and Managed Services for IT Operations.

We’re also the proud creators of Tech in Motion and the Timmy Awards, our North American community platform, events series and award program that connects over 250,000 tech enthusiasts to meet, learn, and innovate.

About_Company_OneAbout_Company_Two
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Easy Apply

Full-time

$80000 - $120000

Remote

Today

Easy Apply

Full-time

$170000 - $200000

Remote

Today

Easy Apply

Full-time

$180000 - $250000

Remote or Weehawken Township, New Jersey

Today

Easy Apply

Full-time

$175000 - $200000

Search all similar jobs