Senior DevOps Engineer – HPC / EDA / SLURM

Remote • Posted 4 hours ago • Updated 4 hours ago
Contract W2
6 Months
No Travel Required
Remote
65+
Company Branding Image
Fitment

Dice Job Match Score™

🫥 Flibbertigibetting...

Job Details

Skills

  • ITSM / Documentation
  • DevOps
  • Monitoring
  • Storage
  • Virtualization
  • Identity
  • Provisioning
  • Linux
  • Automation
  • HPC / EDA

Summary

Senior DevOps Engineer – HPC / EDA / SLURM

Location: Remote
Duration: 3 Months
Experience: 5+ Years
Employment Type: Contract

About the Role

We are seeking a Senior DevOps Engineer to support a High Performance Computing (HPC) and Electronic Design Automation (EDA) infrastructure team on a 3-month contract engagement. This role will work closely with the IT Datacenter organization and requires a highly experienced engineer who can operate independently with minimal ramp-up time.

The ideal candidate will have strong hands-on expertise in Linux HPC environments, SLURM, Ansible, infrastructure automation, enterprise identity management, storage, and datacenter migrations.

Key Responsibilities

HPC / EDA Platform Operations

  • Administer and support SLURM-based HPC compute environments, including partition configuration and migration planning.
  • Plan and execute HPC/EDA compute and storage infrastructure migrations across datacenters.
  • Develop migration strategies, assess risks and dependencies, and define operational trade-offs.
  • Create formal Method of Procedure (MOP) documents, implementation plans, and operational runbooks.
  • Coordinate with EDA/SPE, storage, identity, and infrastructure teams to execute platform changes.
  • Validate storage volumes, application access, authentication, and service continuity following migrations.
  • Define HPC storage service tiers and gather performance and capacity requirements for EDA workloads.

Linux Systems Engineering

  • Administer production SUSE Linux Enterprise Server (SLES) 12 and SLES 15 environments.
  • Build and maintain custom Linux OS images and installation media using Kiwi NG.
  • Support bare-metal provisioning using RackN / Digital Rebar Provision and PXE-based deployment.
  • Provision and configure VMware vSphere virtual machines supporting HPC services.
  • Troubleshoot Linux services, system daemons, operational scripts, and production issues.

Automation & Infrastructure as Code

  • Develop and maintain Ansible playbooks and roles for Linux configuration, authentication, and platform automation.
  • Ensure compatibility across multiple SLES versions.
  • Manage infrastructure code through Git/GitHub and maintain artifacts through Artifactory.
  • Participate in pull requests, code reviews, and inner-source infrastructure development.
  • Manage production changes through ServiceNow change-management processes.

Identity & Access Management

  • Configure and integrate enterprise identity technologies including:
    • Okta
    • Active Directory
    • LDAP
    • NIS
    • VAS
    • SSSD
  • Audit and reconcile Linux user/group information, including UID/GID mappings.
  • Troubleshoot authentication and access issues across HPC compute and storage environments.
  • Extend SSSD-based corporate authentication to new compute environments.

Monitoring & Operational Support

  • Evaluate and implement log-management solutions, including Splunk integration.
  • Troubleshoot production services such as VNC, NIS, AutoFS, Zabbix, and related Linux services.
  • Develop technical documentation, architecture diagrams, implementation guides, and end-user documentation using Confluence.

Required Qualifications

  • 5+ years of experience in DevOps, Platform Engineering, Linux Systems Engineering, or a related role.
  • Hands-on experience administering HPC clusters using SLURM or comparable workload managers.
  • Experience supporting EDA, scientific computing, semiconductor, or high-performance engineering environments.
  • Strong production experience with Ansible, including playbook and role development.
  • Experience with bare-metal provisioning tools such as RackN / Digital Rebar, Cobbler, or equivalent.
  • Proven experience planning and executing datacenter or infrastructure migrations.
  • Strong understanding of enterprise Linux authentication and identity technologies, including SSSD, LDAP, Active Directory, NIS, or Okta.
  • Experience with NetApp or comparable enterprise storage platforms in HPC environments.
  • Ability to create detailed MOPs, runbooks, architecture diagrams, and technical documentation.
  • Strong communication skills and ability to coordinate across multiple technical teams.

Required Technical Skills

CategoryRequired Skills
HPC / EDASLURM, HPC compute & storage, EDA infrastructure, datacenter migration
LinuxSLES 12/15, Linux services, ESXi 8.0, Kiwi NG, dracut
AutomationAnsible, YAML, Python, Bash, Perl
ProvisioningRackN / Digital Rebar Provision, PXE
VirtualizationVMware vSphere
IdentitySSSD, Okta, AD, LDAP, NIS, VAS
StorageNetApp SVM, NFS, AutoFS, RootSquash, IOPS/capacity planning
DevOpsGit, GitHub, Artifactory
MonitoringSplunk, Zabbix, Linux log management
ITSM / DocumentationServiceNow, Jira, Confluence, MOPs, technical diagrams

Preferred Qualifications

  • Enterprise experience with SLES 12 and SLES 15.
  • Hands-on experience with RackN / Digital Rebar Provision.
  • Experience creating custom OS images using Kiwi NG or similar tools.
  • VMware vSphere experience supporting HPC infrastructure.
  • Experience migrating large configuration artifacts and binaries to Artifactory.
  • Background in semiconductor, storage, EDA, or high-tech manufacturing IT environments.
  • Experience with enterprise-scale infrastructure modernization and migration programs.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91103492
  • Position Id: 9050815
  • Posted 4 hours ago

Company Info

About Netsynk

Netsynk is a pioneer in the staffing world.' Our mission has been to serve our clients, employees, partners and community by doing it right. People first is our motto. Finding the right candidate for a client or finding a right job for a candidate can sometimes seem a Herculean task. Our goal is to make this task, just not a smooth and easy, but also an enriching and fun experience for them both. Our experience and helpful network helps place talented professionals in organizations across the country, making the experience for both the client and the consultant an enriching one. At Netsynk, the job is where the dream is built. Irrespective of who you are, a recent graduate or high level professional looking for your next good break or even a company looking for help in staffing; we can definitely help. Vision Statement Be the sought after staffing services company for our clients and consultants Mission Statement Serve our clients, employees, partners and community by doing it right.

Contact the job poster
TJ

Tanya Jaiswal

Recruiter @ Netsynk
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs