Senior HPC DevOps Engineer

Rancho Cordova, CA, US • Posted 1 day ago • Updated 29 minutes ago
Contract Corp To Corp
Contract Independent
Contract W2
12 Months
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Apache Cordova
  • IaaS
  • Identity Management
  • VNC
  • Splunk
  • Management
  • Confluence
  • Documentation
  • NAND
  • Migration
  • Semiconductors
  • Storage
  • Manufacturing
  • DevOps
  • Systems Engineering
  • Data Analysis
  • EDA
  • Computational Science
  • Terraform
  • SLES
  • Ubuntu
  • Microsoft Azure
  • Cloud Computing
  • Linux
  • Authentication
  • LDAP
  • Active Directory
  • NetApp
  • NFS
  • Enterprise Storage
  • HPC
  • Scripting
  • Python
  • Bash
  • YAML
  • Ansible
  • Git
  • GitHub
  • ServiceNow
  • Change Management
  • Technical Writing
  • Communication
  • Oracle UCM
  • OM
  • WebKit
  • SANS
  • IMG

Summary

Senior DevOps Engineer HPC / EDA / SLURM / Azure

Location: Remote Rancho Cordova, CA
Duration: 12 months

Job Description

Role Summary:
Client is seeking a Senior DevOps Engineer to support its HPC and EDA cloud infrastructure. The role requires strong hands-on experience with Linux HPC environments, SLURM, Ansible, Terraform, Azure, identity management, and enterprise storage. The candidate should be able to work independently and support production infrastructure with minimal ramp-up.

Key Responsibilities

  • Administer SLURM-based HPC clusters, compute/storage environments, and EDA infrastructure.
  • Develop and maintain Ansible playbooks/roles and Terraform automation.
  • Support Linux environments including SLES 15 and Ubuntu.
  • Manage enterprise authentication using SSSD, LDAP, Active Directory, and Okta.
  • Support NetApp/NFS, AutoFS, RootSquash, storage capacity and IOPS planning.
  • Manage Azure EDA user environments, including ThinLinc/VNC.
  • Use Git/GitHub, Artifactory and participate in code reviews/PRs.
  • Implement monitoring/logging using Splunk and Datadog.
  • Troubleshoot production Linux services and operational issues.
  • Manage infrastructure changes through ServiceNow.
  • Create MOPs, runbooks, architecture diagrams, implementation guides, and Confluence documentation.
  • Coordinate infrastructure changes with EDA, NAND, storage, and IAM teams.

Must-Haves

  • 5+ years in DevOps, Platform Engineering, or Linux Systems Engineering.
  • Strong hands-on HPC cluster administration experience.
  • SLURM or equivalent workload manager experience.
  • Ansible + Terraform production automation.
  • Strong Linux/SLES/Ubuntu experience.
  • Experience with Azure cloud compute.
  • SSSD, LDAP, AD, Okta / enterprise Linux authentication.
  • NetApp/NFS or comparable enterprise HPC storage.
  • Experience supporting EDA/scientific computing environments.
  • Strong Python/Bash/YAML scripting skills.
  • Experience with Git/GitHub and ServiceNow.
  • Strong technical documentation and cross-team communication skills.

Nice-to-Haves

  • SLES 12/15 enterprise experience.
  • Artifactory migration experience.
  • Semiconductor, storage, or high-tech manufacturing IT background.

Must-Haves

1. 5+ years of DevOps / Platform Engineering / Linux Systems Engineering experience.

2. Strong HPC cluster administration experience.

3. Hands-on SLURM workload manager experience.

4. Experience supporting EDA or scientific computing environments.

5. Strong Ansible experience - playbooks, roles, and production automation.

6. Strong Terraform / Infrastructure as Code experience.

7. Advanced Linux skills, preferably SLES 15 and/or Ubuntu.

8. Hands-on Azure cloud compute experience.

9. Enterprise Linux identity/authentication experience with SSSD, LDAP, Active Directory, and/or Okta.

10. NetApp / NFS or comparable enterprise storage experience in HPC environments.

11. Strong scripting skills with Python and Bash; YAML/Ansible.

12. Git/GitHub experience, including PRs/code reviews.

13. Experience with ServiceNow change management.

14. Ability to create MOPs, runbooks, architecture diagrams, and technical documentation.

15. Strong communication and ability to work independently with multiple infrastructure/engineering teams.

Navnish kumar

Sr. IT Technical Recruiter

Stellent IT Phone:

Email: navnish
Gtalk: navnishom

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91022079
  • Position Id: 2026-51859
  • Posted 1 day ago
Create job alert
Never miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Rancho Cordova, California

•

Today

Easy Apply

Contract

DOE

Remote or Hybrid in Rancho Cordova, California

•

2d ago

Easy Apply

Contract

Depends on Experience

Remote

•

2d ago

Easy Apply

Contract

65 - 70

Utah

•

Today

Full-time

USD 106,300.00 - 206,200.00 per year

Search all similar jobs