Linux Administrator- AI,HPC Infrastructure


Cloud Destinations LLC
Dice Job Match Score™
🔗 Matching skills to job...
Job Details
Skills
- Ansible
- Artificial Intelligence
- BMC
- CPU
- GPU
- HPC
- Firmware
- Linux Administration
- Linux
- PXE
- Python
- Red Hat Enterprise Linux
- Stacks Blockchain
- iPXE
- Ubuntu
- Shell Scripting
- Network
- BIOS
Summary
- Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.
- Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.
- Diagnose failures spanning BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
- Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
- Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
- Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
- Automate repeatable administration and remediation tasks with Bash and Python.
- Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.
- 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
- Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
- Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
- Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
- Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
- Strong shell scripting and Python-based automation capability.
- Working knowledge of storage and network dependencies affecting Linux host health.
- Ability to operate independently in ambiguous, high-severity production situations.
- Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
- Familiarity with GPU telemetry and health monitoring tooling (e.g., DCGM) and high-performance fabric technologies, including telemetry-driven health analysis.
- Experience supporting validation labs or pre-production cluster certification.
- Dice Id: 91097117
- Position Id: 9056880
- Posted 1 hour ago
Company Info
One of the leading US-based staffing and IT consulting partner. Experience exceptional service and top-tier talent across industries. Count on us for staffing solutions that cater to the unique demands of the American market.
Our experienced recruiters ensure a seamless fit within your team, accelerating success. But we go beyond staffing and empower employees with fully sponsored certification programs, keeping them ahead. Experience comprehensive benefits including health, wellness coverage, dental insurance, vision insurance, as well as flexible hours, remote work options, and a robust 401K plan to ensure a secure future at the companies we represent.
At Cloud Destinations, we bring industry expertise and a passion for excellence. From Enterprise Cloud Strategy to Managed Infrastructure Services, Digital Transformation, BI & Data Analytics, Security, Data Engineering, and more, we navigate the IT landscape with finesse. Choose us as your trusted partner, witness transformative talent and exceptional service. Let's unlock new possibilities and drive your success in the dynamic world of IT together.

Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs