Core Platform Engineer/ SRE (incident management , Storage, Network, GPU)

Sunnyvale, CA, US • Posted 11 hours ago • Updated 11 hours ago
Contract W2
Contract Independent
Contract Corp To Corp
12 Months
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Ansible
  • GPU
  • Incident Management
  • Linux Kernel
  • Network Programming
  • Python
  • Physical Layer
  • SAN
  • Sourcing
  • Terraform
  • SLOs/SLIs

Summary

Role 1 — Core Platform Engineer, Level 1
The engineering first line of defense: incident response, triage, reliability, and automation across the full infrastructure stack.
Day-to-day:
  • Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence.
  • Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies.
  • Participates in team stand-ups on projects, incidents, and daily priorities.
  • Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring.
  • Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment.
  • Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset.
Must have:
  • Architecture, design patterns, reliability, and scaling of new and existing systems.
  • Incident command experience — driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through.
  • Observability built from the ground up — defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do.
  • Linux kernel internals — scheduler, memory allocation, driver subsystems.
  • High-quality code in at least one language (Python, Go, or similar).
  • System-level debugging — kdump, kernel panic analysis.
  • IaC (Ansible, Terraform, Kubernetes) and CI/CD (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.
  • TCP/IP and network programming.
  • Distributed storage systems — object, block, and/or file storage paradigms.
  • Strong communication skills.
Sourcing note: This is a deep SRE profile, not a pure generalist. The kernel-internals and system-level debugging bar is real and higher than a typical "L1" label implies — screen for genuine engineering depth, not helpdesk/NOC-tier breadth.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91165977
  • Position Id: 9076791
  • Posted 11 hours ago

Company Info

About Balin Technologies LLC

Balin Technologies, headquartered in Cumming, GA, is one of the leading IT consulting firms founded by industry experts with extensive experience in IT consulting services, Talent Acquisition, and SOW outlining scope, timeline, cost, and other aspects between two parties. Our priority is customer satisfaction, the cornerstone of our success.

We provide end-to-end IT consulting services, from requisition to candidate onboarding, across various industry verticals. Our rigorous screening, interviewing, and recruiting processes ensure the right fit for contract, contract-to-hire, and permanent placements, catering to clients of all sizes.

About_Company_OneAbout_Company_Two
Contact the job poster
Phani Kishore

Phani Kishore

Recruiter @ Balin Technologies LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

San Jose, California

Yesterday

Easy Apply

Contract, Third Party

Depends on Experience

Sunnyvale, California

Today

Easy Apply

Contract, Third Party

Depends on Experience

Search all similar jobs