Role :- Senior Linux Administration
Location :- Santa Clara, CA(Onsite)
Duration: Long Term Contract
AI and HPC Infrastructure
Description
Engagement Summary
The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.
WHAT THIS CANDIDATE WILL BE DOING
· Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.
· Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.
· Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
· Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
· Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
· Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
· Automate repeatable administration and remediation tasks with Bash and Python.
· Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.
WHAT WE NEED TO SEE
· 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
· Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
· Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
· Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
· Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
· Strong shell scripting and Python-based automation capability.
· Working knowledge of storage and network dependencies affecting Linux host health.
· Ability to operate independently in ambiguous, high-severity production situations.
PREFERRED EXPERIENCE
· Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
· Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.
Experience supporting validation labs or pre-production cluster certification