Senior Production Engineer


Neural Strategic Solutions, Inc.
Dice Job Match Score™
⭐ Evaluating experience...
Job Details
Skills
- GPU
- Hardware
- SDN
- OVN
- OVS
- Storage
- Lightbits
- VAST
- Pure Storage
- Root Cause Analysis (RCA)
- Linux Kernel Internals
- Kernel Panic Analysis
- kdump
Summary
Senior Production Engineer
6 Month CTH+ extension
Onsite in Sunnyvale, CA or SFO, CA
JOB DESCRIPTION
Key priorities in candidate evaluation:
Strong hands-on experience in at least one of these areas:
- GPU-based hardware (management, troubleshooting, at-scale operations)
- SDN (OVN, OVS)
- Storage (Lightbits, VAST, Pure Storage)
Candidates with depth across all three will be rare — what matters is genuine expertise in at least one of these domains rather than surface-level familiarity across all of them.
What You''ll Be Working On:
- As a CORE PE at Crusoe, you will engage in incident response drills, post-mortems, and root cause analysis sessions to learn from past issues and prevent future ones.
- Each morning starts with a structured review of overnight alerts and system performance metrics - identifying any anomalies, triaging what needs attention
- You will collaborate with your team in a morning stand-up meeting to discuss ongoing projects, recent incidents, and priorities for the day.
- Your tasks will include automating routine processes, analyzing system logs, and developing tools to enhance our monitoring capabilities.
- You''ll spend part of your day working closely with software engineers, advising on best practices for resilient code and reviewing changes before deployment
- Throughout the day, your focus is on maintaining high SLIs and SLOs, ensuring that our infrastructure remains robust and reliable for our customers.
- By day''s end, you will document your work, share insights with your team, and plan for the next day''s challenges, always with a customer-centric mindset.
What You’ll Bring to the Team:
- Strong experience with architecture, design patterns, reliability and scaling of new and current systems
- Experience leading and commanding incidents, including driving root cause analysis, coordinating cross-functional teams, and ensuring follow-through on corrective actions
- Experience building observability from the ground up — defining SLOs/SLIs, closing monitoring gaps, and implementing alerting strategies that catch failures before customers do
- Proficiency in Linux kernel internals, with exposure to scheduler, memory allocation, and driver subsystem
- Experience writing high quality code with at least one programming language (Python, Go, or similar)
- Experience with system-level debugging, including kdump, and kernel panic analysis.
- Proficiency in Infrastructure as Code tooling (Ansible, Terraform, Kubernetes) and CI/CD practices (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure.
- Experience with TCP/IP and network programming
- Experience with distributed storage systems and understanding of one or more of object, block, and file storage paradigms.
- Hardware and GPU troubleshooting experience (nice to have)
- Exposure to OVN/OVS-based networking stack (nice to have)
- Strong communication skills
- Dice Id: 91166296
- Position Id: 9045711
- Posted 1 day ago
Company Info
About Neural Strategic Solutions, Inc.
Neural Strategic Solutions Inc. is a professional staffing and solutions firm providing flexible and permanent staffing solutions in various IT Domains. Within our staffing offering, we provide temporary, temporary-to-hire, direct placement services for individual as well as full team placements.
Neuralss Inc. has embarked on a journey to deliver exceptional IT consultancy services. Our team is dedicated to providing innovative solutions and exceeding client expectations. We take pride in our unique approach and commitment to delivering quality results that set us apart from the competition.
Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs