Job Description: Senior DevOps Engineer – HPC / EDA
Job Title: Senior DevOps Engineer – HPC / EDA
Contract Duration: 12 Months
Department: IT Datacenter (ITDC)
Work Location: Rancho Cardova, CA (Hybrid/Remote)
Position Summary
We are seeking an experienced Senior DevOps Engineer to support and maintain High-Performance Computing (HPC) and Electronic Design Automation (EDA) infrastructure. The ideal candidate will have strong hands-on experience in Linux systems administration, SLURM workload management, infrastructure automation, Terraform, Ansible, Azure cloud environments, enterprise identity management, and storage administration.
The candidate will work closely with infrastructure, EDA, storage, and Identity and Access Management (IDAM) teams to deliver reliable, secure, and scalable HPC platform operations. This role requires a senior-level engineer who can independently manage production infrastructure changes, troubleshoot complex technical issues, develop automation, and maintain comprehensive technical documentation.
Key Responsibilities
1. HPC / EDA Platform Operations
Administer and support SLURM-based HPC compute environments, including workload management, partition configuration, and infrastructure migration planning.
Plan and execute infrastructure changes and migrations using Terraform and Ansible.
Support Azure-based EDA user environments, including ThinLinc/VNC access and related services.
Coordinate with EDA, TD NAND, storage, and IDAM teams to implement infrastructure changes and platform upgrades.
Develop and maintain formal Methods of Procedure (MOPs), operational runbooks, and service cutover documentation.
2. Automation & Infrastructure as Code
Develop, maintain, and enhance Ansible playbooks and roles for Linux provisioning, authentication, system configuration, and platform administration.
Ensure compatibility of Ansible playbooks across multiple versions and SLES 15 environments.
Use Terraform to automate infrastructure provisioning and support configuration management and migration activities.
Manage infrastructure code through Git and GitHub, including pull requests, code reviews, and internal repository contributions.
Support artifact and binary management using Artifactory.
Implement production changes through established change management processes using ServiceNow.
3. Identity & Access Management
Configure and troubleshoot enterprise authentication and identity integration for Linux and HPC environments.
Work with SSSD, LDAP, Active Directory, and Okta to support centralized authentication and access management.
Audit and reconcile Linux user and group identity information, including UID/GID consistency across multiple directory and authentication domains.
Validate authentication, authorization, and access behavior across compute and storage environments.
Extend SSSD-based corporate authentication to new compute environments and automate configurations using Ansible.
4. Monitoring, Logging & Operational Readiness
Evaluate and implement log management solutions for HPC systems, including potential Splunk integration.
Monitor, troubleshoot, and resolve production Linux service issues involving ThinLinc/VNC, AutoFS, Datadog, and related infrastructure services.
Develop and maintain operational scripts using Python, Bash, and Perl as required.
Support production readiness assessments, infrastructure validation, and operational improvement initiatives.
Create and maintain technical documentation, architecture diagrams, implementation guides, and end-user instructions in Confluence.
5. Enterprise Storage & Filesystem Administration
Support enterprise storage platforms, including NetApp Storage Virtual Machines (SVMs) and comparable storage solutions.
Work with NFS, AutoFS, RootSquash, and Linux filesystem configurations.
Assist with storage tier design, capacity planning, and IOPS performance considerations.
Coordinate storage-related changes with infrastructure and HPC teams to ensure platform reliability and performance.
Required Skills & Qualifications
5+ years of experience in DevOps, Platform Engineering, Linux Systems Engineering, or a related infrastructure role.
Hands-on experience administering HPC clusters using SLURM or an equivalent workload manager.
Experience supporting EDA, scientific computing, or high-performance computing environments.
Strong hands-on expertise in Ansible, including playbook and role development, automation, and configuration management.
Experience with Terraform and Infrastructure as Code (IaC).
Strong Linux administration skills, preferably with SUSE Linux Enterprise Server (SLES 15) and Ubuntu.
Experience with Microsoft Azure cloud compute environments.
Working knowledge of SSSD, LDAP, Active Directory, and Okta in enterprise Linux environments.
Experience with NetApp or comparable enterprise storage platforms, NFS, AutoFS, and filesystem administration.
Familiarity with Git, GitHub, pull requests, code reviews, and repository management.
Experience with monitoring and logging tools such as Splunk and Datadog.
Scripting skills in Python, Bash, and/or Perl.
Experience using ServiceNow for production change management.
Ability to create MOPs, runbooks, architecture diagrams, and technical implementation documentation.
Strong troubleshooting, analytical, communication, and cross-functional collaboration skills.
Ability to work independently and manage complex infrastructure tasks with minimal supervision.
Preferred Qualifications
Experience with SLES 12 and/or SLES 15 in enterprise environments.
Experience migrating configuration artifacts and binaries to Artifactory.
Background in semiconductor, NAND, storage, or high-tech manufacturing IT environments.
Experience with HPC datacenter migrations and large-scale infrastructure transitions.
Familiarity with enterprise storage performance optimization and capacity planning.
Ideal Candidate Profile
The ideal candidate is a hands-on senior infrastructure engineer with a strong combination of Linux administration, HPC/EDA operations, SLURM, Ansible, Terraform, Azure, enterprise authentication, and storage expertise. The candidate should be comfortable troubleshooting complex production environments, automating repetitive tasks, documenting infrastructure changes, and collaborating with multiple technical teams.