Sr Platform Engineer

San Jose, CA, US • Posted 30+ days ago • Updated 3 days ago
Contract Corp To Corp
Contract W2
Contract Independent
12 Months
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • CPU
  • NVMe
  • GPU
  • Linux
  • HPC
  • Network
  • incident

Summary

Job Title: Core Platform /System Engineer
Location: Sunnyvale, CA ,  San Jose (On-Site)
Duration: Long-Term Contract
Job Description
  • As a Core Platform Engineer, we will serve as the first responder for production incidents,
  • orchestrate incident management, drive reliability improvements, and establish SRE best
  • practices across Compute, Networking, Storage, and GPU infrastructure teams.
  • Candidates should possess hands-on infrastructure experience and sufficient technical
  • depth to identify affected systems, engage the right subject matter experts, and drive
  • incident resolution processes using data and observability signals.
Responsibilities
  • Act as first responder during infrastructure incidents.
  • Lead incident bridges and coordinate cross-functional response efforts.
  • Perform incident triage and identify impacted infrastructure domains.
Gather evidence and telemetry to route incidents to the correct SME team.
  • Drive incident communications and stakeholder updates.
  • Improve reliability processes across platform engineering teams.
  • Define and promote SRE best practices and operational standards. (SLO,SLI)
  • Identify observability gaps and implement improvements.
  • Build automation for incident response workflows.
  • Manage and optimize incident management tooling (e.g., ).
  • Support change management and operational readiness processes.
  • Assist foundation engineering teams in identifying reliability risks and trends.
  • Participate in on-call activities and operational reviews.
Required Qualifications
  • 5 to 10+ years of experience in Site Reliability Engineering, Platform Engineering,
  • Infrastructure Operations, or Systems Engineering.
  • Strong infrastructure troubleshooting experience.
Deep expertise in at least one of the following:
  1. GPU infrastructure
  2. KVM/virtualization
  3. SDN (OVN/OVS)
  4. Storage (Lightbits/Pure Storage)
  • Proven incident management and operational leadership experience.
  • Experience running high-severity production incidents.
  • Strong understanding of observability, monitoring, SLIs, and SLOs.
  • Experience building operational automation.
  • Ability to make data-driven decisions during outages and service disruptions.
  • Preferred Qualifications
  • Experience with  or similar incident management platforms.
  • Kubernetes production operations experience.
  • Cloud-native infrastructure experience.
  • Experience supporting large-scale AI or GPU environments.
  • Strong communication and stakeholder management skills.
  • What Success Looks Like
  • Quickly identifies affected infrastructure domains during incidents.
  • Effectively coordinates SMEs and engineering teams.
  • Reduces incident response and recovery times.
  • Improves observability and operational processes.
  • Establishes reliability standards across Crusoe's infrastructure platform.
  • Ideal Candidate Screening Criteria (For Both Roles)
 
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91165977
  • Position Id: 9047782
  • Posted 30+ days ago

Company Info

About Balin Technologies LLC

Balin Technologies, headquartered in Cumming, GA, is one of the leading IT consulting firms founded by industry experts with extensive experience in IT consulting services, Talent Acquisition, and SOW outlining scope, timeline, cost, and other aspects between two parties. Our priority is customer satisfaction, the cornerstone of our success.

We provide end-to-end IT consulting services, from requisition to candidate onboarding, across various industry verticals. Our rigorous screening, interviewing, and recruiting processes ensure the right fit for contract, contract-to-hire, and permanent placements, catering to clients of all sizes.

About_Company_OneAbout_Company_Two
Contact the job poster
MA

Mahesh Agurla

Recruiter @ Balin Technologies LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Sunnyvale, California

2d ago

Easy Apply

Contract, Third Party

Depends on Experience

Sunnyvale, California

2d ago

Easy Apply

Third Party, Contract

Depends on Experience

Search all similar jobs