Senior Linux Administration

Santa Clara, CA, US • Posted 23 hours ago • Updated 23 hours ago
Contract W2
12 Months
No Travel Required
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Linux Administration
  • HPC
  • Linux
  • AI

Summary

Role :- Senior Linux Administration

Location :- Santa Clara, CA(Onsite)

Duration: Long Term Contract

AI and HPC Infrastructure

Description

Engagement Summary

The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.

WHAT THIS CANDIDATE WILL BE DOING

·                Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.

·                Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.

·                Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.

·                Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.

·                Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.

·                Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.

·                Automate repeatable administration and remediation tasks with Bash and Python.

·                Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.

WHAT WE NEED TO SEE

·                7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.

·                Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.

·                Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.

·                Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.

·                Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.

·                Strong shell scripting and Python-based automation capability.

·                Working knowledge of storage and network dependencies affecting Linux host health.

·                Ability to operate independently in ambiguous, high-severity production situations.

PREFERRED EXPERIENCE

·                Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.

·                Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.

Experience supporting validation labs or pre-production cluster certification

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: prutx001
  • Position Id: 9067836
  • Posted 23 hours ago
Contact the job poster
AS

Anil Swamy Madapath

Recruiter @ Prudent Technologies and Consulting
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Santa Clara, California

Today

Easy Apply

Contract

50 - 60

Santa Clara, California

Today

Easy Apply

Contract

Depends on Experience

Sunnyvale, California

16d ago

Easy Apply

Contract, Third Party

$70 - $75

Fremont, California

7d ago

Easy Apply

Full-time

Depends on Experience

Search all similar jobs