Senior Production Engineer

San Francisco, CA, US • Posted 22 hours ago • Updated 22 hours ago
Contract W2
6 Months
No Travel Required
On-site
$45 - $50/hr
Fitment

Dice Job Match Score™

✨ Finding the perfect fit...

Job Details

Skills

  • Production Engineering
  • SRE
  • large-scale infrastructure operations
  • GPU
  • TCP/IP
  • GitLab

Summary

Job Overview

We are seeking an experienced Senior Production Engineer to support and maintain highly scalable, reliable infrastructure environments. The ideal candidate will have strong hands-on expertise in Linux systems, production operations, incident management, observability, automation, and system-level debugging.

Candidates should have deep expertise in at least one of the following areas:

  • GPU-based hardware, including management, troubleshooting, and large-scale operations
  • Software-Defined Networking (SDN), including OVN and OVS
  • Storage technologies, including Lightbits, VAST, Pure Storage, or distributed storage platforms

Expertise across all three domains is not required. Strong, hands-on depth in at least one domain is preferred over surface-level experience across multiple areas.

Key Responsibilities

  • Monitor production systems, review alerts, analyze system performance, and identify operational anomalies.
  • Lead and participate in incident response, incident drills, post-mortems, and root cause analysis.
  • Coordinate with cross-functional teams during critical production incidents and drive corrective actions.
  • Design and implement observability solutions, including SLIs, SLOs, monitoring, and proactive alerting strategies.
  • Automate routine operational processes and develop tools to improve monitoring and system reliability.
  • Analyze system logs and troubleshoot complex infrastructure and production issues.
  • Collaborate with software engineering teams to improve application resilience and reliability.
  • Review system and infrastructure changes before deployment and recommend reliability best practices.
  • Maintain high availability, scalability, and reliability across production infrastructure.
  • Document operational processes, incident findings, technical solutions, and improvement plans.

Required Qualifications

  • Strong experience with system architecture, design patterns, scalability, reliability, and production operations.
  • Experience leading production incidents, driving root cause analysis, and coordinating cross-functional teams.
  • Strong hands-on experience building observability from the ground up, including defining SLIs/SLOs and implementing monitoring and alerting strategies.
  • Strong knowledge of Linux systems and Linux kernel internals, including scheduling, memory allocation, and driver subsystems.
  • Proficiency in Python, Go, or another systems programming language.
  • Experience with system-level debugging, including kdump and kernel panic analysis.
  • Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, and Kubernetes.
  • Knowledge of CI/CD tools and practices, including GitLab CI, AWX, or similar technologies.
  • Strong understanding of TCP/IP, networking, and network programming.
  • Experience with distributed storage systems and object, block, or file storage architectures.
  • Excellent troubleshooting, analytical, and communication skills.

Preferred Qualifications

  • Hands-on experience with GPU hardware management, troubleshooting, and large-scale GPU infrastructure.
  • Experience with Software-Defined Networking technologies, including OVN and OVS.
  • Experience with Lightbits, VAST, Pure Storage, or similar enterprise storage platforms.
  • Experience supporting large-scale bare-metal or cloud infrastructure.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91125539
  • Position Id: 9047548
  • Posted 22 hours ago
Contact the job poster
Shahbaz Hussain

Shahbaz Hussain

Senior Technical Recruiter @ IntelliTask LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

San Francisco, California

Today

Full-time

USD 225,000.00 - 275,000.00 per year

San Francisco, California

Today

Full-time

USD 200,000.00 - 280,000.00 per year

Sunnyvale, California

2d ago

Easy Apply

Contract

70 - 75

Sunnyvale, California

Today

Easy Apply

Contract, Third Party

Depends on Experience

Search all similar jobs