GPU Infrastructure Site Reliability Engineer(On-site, L2- Face to Face)

Sunnyvale, CA, US • Posted 12 hours ago • Updated 12 hours ago
Contract Corp To Corp
Contract Independent
Contract W2
12 Months
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

⏳ Almost there, hang tight...

Job Details

Skills

  • GPU
  • Embedded Platform Engineering
  • Linux
  • Infrastructure
  • Site Reliability Engineering (SRE)
  • OVS
  • OCS
  • Lightbits Storage
  • Pure Storage
  • Compute (GKN)
  • Incident Management
  • Alert Monitoring
  • Production Support
  • Troubleshooting
  • Root Cause Analysis
  • Infrastructure Operations
  • Networking
  • Storage
  • Compute Infrastructure
  • Linux Administration

Summary

Job Title: Site Reliability Engineer (SRE)

Location: Sunnyvale, CA (On-site)
Duration: Long-Term Contract

Job Description

We are seeking a highly motivated Site Reliability Engineer (SRE) to support mission-critical AI and GPU infrastructure in a high-performance production environment. The ideal candidate will have experience supporting GPU platforms, embedded infrastructure, compute, networking, and storage systems while ensuring high availability, reliability, and operational excellence.

As an SRE, you will monitor production environments, troubleshoot infrastructure issues, respond to incidents, and collaborate with platform, hardware, and engineering teams to maintain scalable and reliable infrastructure.

Key Responsibilities

  • Monitor and maintain production infrastructure supporting GPU and embedded platforms.
  • Investigate and resolve infrastructure incidents, system alerts, and performance issues.
  • Support GPU servers, compute infrastructure, storage systems, and networking components.
  • Perform troubleshooting across Linux systems, hardware, networking, storage, and platform services.
  • Work closely with Platform Engineering, Infrastructure, Network, and Hardware teams to ensure platform reliability.
  • Support deployment, provisioning, configuration, and maintenance of infrastructure components.
  • Monitor infrastructure health and proactively identify reliability and performance issues.
  • Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
  • Participate in infrastructure upgrades, maintenance activities, and production rollouts.
  • Create and maintain operational documentation, runbooks, and standard operating procedures.
  • Participate in on-call rotation and provide production support for critical infrastructure.

Required Skills

  • Strong experience in Site Reliability Engineering (SRE) or Infrastructure Engineering.
  • Hands-on experience supporting GPU-based infrastructure.
  • Experience with Embedded Platform Engineering environments.
  • Strong Linux administration and troubleshooting skills.
  • Good understanding of compute infrastructure.
  • Experience with enterprise infrastructure and production operations.
  • Knowledge of networking concepts with experience in OVS (Open vSwitch) and OCS.
  • Experience with enterprise storage solutions such as Lightbits and Pure Storage.
  • Understanding of GKN Compute environments or similar compute platforms.
  • Experience in infrastructure monitoring, incident management, and alert handling.
  • Strong troubleshooting and root cause analysis skills.
  • Excellent communication and collaboration skills.

Preferred Skills

  • Experience with Kubernetes or container platforms.
  • Knowledge of cloud platforms (AWS, Azure, or Google Cloud Platform).
  • Familiarity with automation using Bash, Python, or Ansible.
  • Experience with monitoring tools such as Prometheus, Grafana, Datadog, or Splunk.
  • Knowledge of CI/CD and Infrastructure as Code (Terraform, Ansible).

Mandatory Skills

  • GPU Infrastructure
  • Embedded Platform Engineering
  • Linux Administration
  • Infrastructure Operations
  • Compute (GKN or similar)
  • OVS (Open vSwitch)
  • OCS Networking
  • Lightbits Storage
  • Pure Storage
  • Incident Management
  • Alert Monitoring
  • Root Cause Analysis
  • Production Support
  • Troubleshooting
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91165977
  • Position Id: 9047795
  • Posted 12 hours ago

Company Info

About Balin Technologies LLC

Balin Technologies, headquartered in Cumming, GA, is one of the leading IT consulting firms founded by industry experts with extensive experience in IT consulting services, Talent Acquisition, and SOW outlining scope, timeline, cost, and other aspects between two parties. Our priority is customer satisfaction, the cornerstone of our success.

We provide end-to-end IT consulting services, from requisition to candidate onboarding, across various industry verticals. Our rigorous screening, interviewing, and recruiting processes ensure the right fit for contract, contract-to-hire, and permanent placements, catering to clients of all sizes.

About_Company_OneAbout_Company_Two
Contact the job poster
Phani Kishore

Phani Kishore

Recruiter @ Balin Technologies LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

San Jose, California

Today

Easy Apply

Contract

Depends on Experience

Sunnyvale, California

Today

Easy Apply

Full-time

70,000 - 80,000

Sunnyvale, California

Yesterday

Easy Apply

Full-time, Third Party

900000

Search all similar jobs