Staff HPC Infrastructure Engineer

• Posted 11 hours ago • Updated 11 hours ago
Full Time
USD $155,700.00 - 214,150.00 per year
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Analytics
  • LinkedIn
  • Twitter
  • Facebook
  • Backbone.js
  • Genomics
  • Bioinformatics
  • Research and Development
  • Data Processing
  • Data Integration
  • Teamwork
  • Agile
  • SLA
  • Research
  • Shell
  • Scripting
  • Python
  • Reliability Engineering
  • Interfaces
  • Accessibility
  • Mentorship
  • Repair
  • Quality Assurance
  • DevOps
  • Performance Tuning
  • Ethernet
  • Enterprise Networks
  • IaaS
  • Partnership
  • Storage
  • MSP
  • Computer Science
  • Unix Administration
  • TCP/IP
  • Ansible
  • InfiniBand
  • Remote Direct Memory Access
  • Data Storage
  • High Performance Computing
  • HPC
  • Amazon Web Services
  • Google Cloud
  • Google Cloud Platform
  • Microsoft Azure
  • System Administration
  • Technical Writing
  • Cisco
  • Network
  • Cisco Certifications
  • Arista
  • Computer Networking
  • IBM
  • IBM GPFS
  • Linux
  • Red Hat Linux
  • Debian Linux
  • SUSE Linux
  • Cloud Computing
  • File Systems
  • Docker
  • Kubernetes
  • HIPAA
  • Sarbanes-Oxley
  • Fluency
  • Artificial Intelligence
  • Management
  • Collaboration
  • Science
  • System Integration Testing
  • Recruiting
  • Screening
  • Law
  • SAP BASIS
  • Privacy

Summary

Company Description

Guardant Health is a leading precision oncology company focused on guarding wellness and giving every person more time free from cancer. Founded in 2012, Guardant is transforming patient care and accelerating new cancer therapies by providing critical insights into what drives disease through its advanced blood and tissue tests, real-world data and AI analytics. Guardant tests help improve outcomes across all stages of care, including screening to find cancer early, monitoring for recurrence in early-stage cancer, and treatment selection for patients with advanced cancer. For more information, visit guardanthealth.com and follow the company on LinkedIn, X (Twitter) and Facebook.

Guardant's HPC team builds and operates the computational technology backbone of the company. This includes scalable data storage that holds petabytes of genomics data, high-performance compute clusters running a custom bioinformatics pipeline in production and R&D environments, and the software infrastructure that hosts an ecosystem of services for internal data processing and external data integration.

The HPC engineering team is looking for a Staff-level engineer with broad, all-round HPC competency and specialized depth in one or more of the following: Red Hat-family OS management, networking, storage, Kubernetes, or Slurm. Experience applying these skills both on-premise and in the cloud is a plus, as we continue to evolve our HPC footprint. This role carries technical depth and cross-team representation for HPC compute, networking, and storage-integration initiatives, partnering closely with our engineering team and our managed service provider (MSP) as we scale operations. To support Guardant Health's fast growth over the next few years, we need a strong technical engineer who can help maintain and grow the HPC infrastructure through this expansion while working closely with corporate IT, SQA, and DevOps/SRE teams. Strong teamwork and the ability to manage multiple in-flight, cross-functional projects at once are essential for success in this role.

About the Role

You enjoy an agile, very fast paced and highly technical environment. You are a self-driven, accomplished technologist who strives to continually improve your skills as the HPC landscape evolves and the computational infrastructure scales. You are dedicated to engineering excellence yet pragmatic and flexible. You have the ability to maintain the day-to-day support SLA while running various key projects that move the business forward. You are comfortable being the deepest technical voice in the room on your declared specialism, even among more senior colleagues, while operating as a peer to the other engineers on the team.

Essential Duties and Responsibilities

Broad HPC Skills (all-round)

Manage multiple HPC clusters and cluster file systems

Integrate cloud bursting as part of the HPC abstraction work

Research, develop, and implement the next generation HPC solutions

Troubleshoot the production system stack down to source code level, e.g shell scripts, Python, and others

Maintain, monitor, and support the infrastructure environment and/or facilities

Use and maintain enhanced production monitoring and addition capability

Support improvements for increased system reliability and performance

Support multiple systems or applications of medium to high complexity complexity defined by size, technology used, and system feeds and interfaces) with multiple concurrent users, ensuring control, integrity, and accessibility

Support systems at remote locations, including internationally

Mentor junior engineers on HPC best practices

Work with offsite consultants to maintain the infrastructure

Work with vendors to troubleshoot, upgrade, and repair systems as needed

Represent HPC infrastructure networking and storage-integration topics in cross-functional planning with networking, SQA, DevOps/SRE, and the MSP

Set up and ownership supporting XDMoD instances for HPC metric and monitoring

Participate in a 24/7 on-call rotation

Networking

Act as the technical peer for HPC networking and interconnect initiatives with the dedicated networking engineer

Collaborate on design, performance tuning, and troubleshooting of HPC Ethernet

Work with enterprise networking on integration of HPC systems with the bandwidth-on-demand system that connects our sites and cloud infrastructure

Work with the networking infrastructure team to manage and optimize connectivity to and from HPC systems and global locations

Storage

Act as the technical peer for the architecture and integration strategy for. HPC storage in partnership with the dedicated storage engineer and MSP

Serve as a technical point of contact for the MSP storage relationship and help define and evolve SLAs, validate delivery, and escalate technical issues

Support the transition of day-to-day storage operations to the MSP without loss of performance or reliability

Required Qualifications
  • Bachelor's degree in Computer Science or a related field with 8-12 years of relevant experience; Master's degree with 6-8 years of relevant experience; or PhD with 3-5 years of relevant experience
  • Strong experience in systems and/or infrastructure engineering, including Linux/Unix administration and TCP/IP networking.
  • Hands-on experience with automation tools, such as Ansible or equivalent technologies.
  • Experience with high-performance networking technologies, such as InfiniBand, RoCE, RDMA, or equivalent, including troubleshooting in production environments.
  • Experience supporting large-scale data storage and high-performance computing (HPC)/compute environments.
  • Experience working with both on-premise and cloud-based infrastructure, such as AWS, Google Cloud Platform (Google Cloud Platform), Azure, or similar environments.
  • Experience developing and supporting software release, operations, and infrastructure automation processes and toolsets.
  • Strong experience creating and maintaining system administration and technical documentation.

Preferred Qualifications
  • Cisco Certified Network Professional (CCNP) certification
  • Experience with Arista and compatible networking, up to and including 400 Gb/s links
  • Experience administering IBM's General Parallel File System (GPFS)
  • Experience administering the Slurm scheduler
  • Experience using Warewulf
  • Linux support and OS management. Red Hat family is a must, but Debian or Suse is nice to have.
  • Experience with cloud bursting technologies
  • Experience with wide area file systems
  • Experience with Docker and Apptainer container technologies
  • Experience with Kubernetes
  • Operating infrastructure compliant with HIPAA and SOX standards

AI & Digital Fluency
  • Demonstrate curiosity, sound judgment, and the ability to critically evaluate and responsibly leverage AI-enabled tools in accordance with company policies, ethical standards, and regulatory requirements to improve the efficiency, effectiveness, and quality of work.

Hybrid Work Model:This section is applicable to onsite employees who are eligible for hybrid work location as specified by management and related policies. Guardant has defined days for in-person/onsite collaboration and work-from-home days for individual-focused time. All U.S. employees who live within 50 miles of a Guardant facility will be required to be onsite on Mondays, Tuesdays, and Thursdays. We have found aligning our scheduled in-office days allows our teams to do the best work and creates the focused thinking time our innovative work requires. At Guardant, our work model has created flexibility for better work-life balance while keeping teams connected to advance our science for our patients.

The annualized base salary ranges for the primary location and any additional locations are listed below. This range does not include benefits or, if applicable, bonus, commission, or equity. Each candidate's compensation offer will be based on multiple factors including, but not limited to, geography, experience, education, job-related skills, job duties, and business need.Primary Location: Remote-USA-CAPrimary Location Base Pay Range: $155,700 - $214,150Other US Location(s) Base Pay Range: $147,100 - $202,300If the role is performed in Colorado, the pay range for this job is: $155,700 - $214,150

Employee may be required to lift routine office supplies and use office equipment. Majority of the work is performed in a desk/office environment; however, there may be exposure to high noise levels, fumes, and biohazard material in the laboratory environment. Ability to sit for extended periods of time.

Guardant Health is committed to providing reasonable accommodations in our hiring processes for candidates with disabilities, long-term conditions, mental health conditions, or sincerely held religious beliefs. If you need support, please reach out to

A background screening including criminal history is required for this role. GH will consider qualified applicants with criminal arrest or conviction histories in a manner consistent with applicable law including but not limited to the LA County Fair Chance Policies and the Fair Chance Act (Gov. Code Section 12952).

Guardant Health is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability.

All your information will be kept confidential according to EEO guidelines.

To learn more about the information collected when you apply for a position at Guardant Health, Inc. and how it is used, please review our Privacy Notice for Job Applicants.

Please visit our career page at:
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90979508
  • Position Id: a5c900ec51a3fff3b8bdd989c0d8d290
  • Posted 11 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Chicago, Illinois

Today

Full-time

USD 150,000.00 - 200,000.00 per year

Rockville, Maryland

Today

Full-time

USD 123,250.00 - 166,750.00 per year

New York, New York

Today

Full-time

USD 150,000.00 - 200,000.00 per year

Bothell, Washington

Today

Full-time

USD 145,920.00 - 209,241.00 per year

Search all similar jobs