RedLine Performance Solutions (RedLine) has been in the HPC solutions engineering services business for 26 years and is consistently determined to keep the bar of excellence quite high for new hires. This enables RedLine to accomplish what other firms cannot and promotes a high level of staff retention. We offer services ranging from full life cycle HPC systems engineering to remote managed services to HPC program analysis.
RedLine is looking for a Senior High Performance Computing (HPC)/Cloud Systems Administrator to join our team. This position will work on various on-premise and cloud-based installations in support of customers who have a wide range of requirements. The Lead Systems Administrator will be an experienced individual with a strong security, Linux, HPC, configuration management, systems automation and networking background.
This position supports customers who deliver mission-critical programs, providing life-changing and economy shaping services. By ensuring the success of their technical programs, you will be directly contributing to efforts that protect the public, drive economy, and improve people's lives.
U.S. Citizenship and the ability to obtain a Public Trust clearance is a requirement to apply. This is a remote position. This full-time position offers a full benefits package including paid time off, 401k match, and health care benefits.
Job Responsibilities:
- Lead a team to administer resources from the operating system and above within on-premise and cloud-based HPC environments. Efforts include, but are not limited to:
- Integration and configuration of compute resources and service nodes
- All software installations on compute resources, services nodes, and parallel filesystems
- Develop/implement system and performance monitoring and benchmarking
- Maintain system documentation
- Evaluate performance impacts of planned operating system changes
- Lead resource optimization and job scheduling software and policies
- Provide technical support to researchers using HPC resources, troubleshoot problems and develop appropriate computational strategies
- Provide emergency support on a 24x7 basis.
- Provide technical leadership and direction for other team members. Maintain team focus on production uptime and model performance.
- Review and present all change management requests to customer management.
- Prepare for and attend the daily 9:00 am (Eastern) Operations call with customer staff and other operational stakeholders.
Requirements:
- Minimum of 10 years RedHat, Rocky, and/or CentOS Linux system administrator experience.
- Technical leadership experience in a large, production environment – leadership for the technical solution and the system administration team.
- Demonstrated ability to configure, deploy and manage major system areas such as batch system, network, data storage, backup system, database system, or distributed computing
- Experience with configuration management tools (e.g., Ansible)
- Ability to work both independently and as part of the team; flexibility in dealing with assignments and in working on several projects simultaneously
- Ability to effectively communicate with people of diverse backgrounds and computer knowledge.
Preferred Skills:
- HPC system administration experience is highly preferred.
- Experience with batch systems (e.g., SLURM)
- Experience managing parallel and cluster file systems (e.g., Lustre)
- Network management experience, including in an HPC context (e.g., InfiniBand)
- Experience with cloud HPC environment (e.g., Google Cloud Platform) is highly preferred.