HPC Infrastructure & Cluster Engineer

Springfield, VA, US • Posted 4 days ago • Updated 9 hours ago
Full Time
On-site
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • HPC
  • Procurement
  • Information Technology
  • IT Service Management
  • Internal Communications
  • IC
  • Integrated Circuit
  • Linux
  • Distribution
  • Machine Learning (ML)
  • Optimization
  • Operating Systems
  • Network Administration
  • Computer Networking
  • Network Design
  • Technology Integration
  • ProVision
  • Red Hat Linux
  • Regulatory Compliance
  • Access Control
  • Security Clearance
  • Continuous Integration
  • Training
  • Military
  • Linux Administration
  • High Performance Computing
  • DoD
  • Security+
  • Customer Engagement
  • SSCP
  • GSEC
  • Cisco Certifications
  • Servers
  • Enterprise Storage
  • InfiniBand
  • Artificial Intelligence
  • Orchestration
  • Kubernetes
  • Writing
  • Scripting
  • Bash
  • Python
  • FOCUS
  • Computer Hardware
  • Network
  • File Systems
  • Storage
  • Management
  • GPU
  • Communication
  • Analytics
  • Systems Engineering
  • Project Management
  • Partnership
  • Cyber Security
  • Collaboration
  • Microsoft Excel
  • Recruiting
  • Database

Summary

Overview

Abile Group has an exciting and challenging opportunity for a HPC Infrastructure & Cluster Engineer on a 10 year contract providing User Facing and Data Center Services supporting an Intelligence Community customer. All the personnel on the team will work together to support innovative design, engineering, procurement, implementation, operations, sustainment and disposal of user facing and data center information technology (IT) services on multiple networks and security domains, at multiple locations worldwide, to support the IC mission.

The right candidate will possess the below skills and qualifications and be ready to handle all responsibilities independently and professionally.

Responsibilities

  • Cluster Administration: Manages the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configures, maintains, and optimizes workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tunes cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administers storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partners with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensures all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.

Qualifications

Clearance Required: TS/SCI with ability to obtain a CI Poly.

Degree and Years of Experience: Bachelor's Degree in a related discipline, or the equivalent combination of education, professional training, or work/military experience.
  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.

Required Certifications:
  • Meet DoD 8570 IAT Level II requirements including one of the following: Security+ CE, CND, SSCP, GSEC, GICSP, CySA+, or CCNA.

Required Skills:
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.

Desired Skills:
  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.

About Abile Group, Inc.

Abile Group, founded in July 2004 to support the Intelligence Community and its contractors across Enterprise Analytics, IT & Systems Engineering, and Program & Project Management, merged with Valiant Solutions in January 2026 - an established provider of cybersecurity technologies and services for Federal Agencies since 2005. Together, this partnership creates a stronger, more integrated cybersecurity organization with expanded opportunities for employees, deeper technical collaboration, and a unified mission. With significant experience serving the Federal Government, we remain dedicated to our employees and clients and seek high-performing professionals who excel at providing guidance, developing solutions, and delivering implementation support that blends industry best practices with client expertise and Abile's broad technical capabilities.

Hiring Statement

Abile is committed to hiring the most qualified and best fit person for the job - always has, always will. Anyone requiring reasonable accommodations should email with requested details. A member of the HR team will respond to your request within 2 business days.

Please review our current job openings and apply for the positions you believe may be a fit. If you are not an immediate fit, we will also keep your resume in our database for future opportunities.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10225783
  • Position Id: 7d440e2076711e1df13c4be01a9796dc
  • Posted 4 days ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Springfield, Virginia

Today

Full-time

USD 140,000.00 - 185,000.00 per year

Springfield, Virginia

Today

Full-time

USD 119,850.00 - 162,150.00 per year

Springfield, Virginia

Today

Full-time

USD 148,000.00 - 179,000.00 per year

Arlington, Virginia

Today

Full-time

USD 99,000.00 - 206,000.00 per year

Search all similar jobs