
Prudent Technologies and Consulting
Santa Clara, California • Yesterday
Easy Apply
Contract
Depends on Experience
592 results (16 new)

Prudent Technologies and Consulting
Santa Clara, California • Yesterday
Easy Apply
Contract
Depends on Experience

Stryde Consulting Services LLC
Santa Clara, California • Today
Easy Apply
Contract
Depends on Experience

Aziro Technologies LLC
Hybrid in San Jose, California • 30+d ago
Easy Apply
Contract
Depends on Experience
Aziro Technologies LLC
Hybrid in San Jose, California • 16d ago
Easy Apply
Contract, Third Party
Depends on Experience








InfiCare
Hybrid in Sunnyvale, California • 6d ago
Easy Apply
Contract, Third Party
Depends on Experience


UnitedHealth Group
Remote or Eden Prairie, Minnesota • Today
Full-time
USD 134,600.00 - 230,800.00 per year



Stryde Consulting Services LLC
Santa Clara, California • Yesterday
Easy Apply
Contract
Depends on Experience








Role :- Senior Linux Administration
Location :- Santa Clara, CA(Onsite)
Duration: Long Term Contract
AI and HPC Infrastructure
Description
Engagement Summary
The Candidate will provide senior Linux administration services across AI and HPC environments supporting GPU clusters, high-performance storage, and data center network-connected compute infrastructure. This role is intended for a hands-on operator who can stabilize production systems, resolve complex node-level failures, and improve fleet reliability at scale.
WHAT THIS CANDIDATE WILL BE DOING
· Administer large-scale Linux environments supporting AI training, inference, and HPC workloads.
· Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets.
· Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID, NVMe, systemd, and GPU driver stacks.
· Build and maintain golden images, provisioning pipelines, configuration baselines, and post-deployment validation procedures.
· Partner with network, platform, storage, and validation teams to isolate cross-domain failures affecting cluster readiness or job execution.
· Investigate performance anomalies involving CPU, memory, NUMA, I/O, interrupts, process scheduling, and kernel tuning.
· Automate repeatable administration and remediation tasks with Bash and Python.
· Produce clear runbooks, failure signatures, and escalation criteria for recurring operational issues.
WHAT WE NEED TO SEE
· 7+ years delivering Linux administration in data center, cloud, AI, or HPC environments.
· Deep expertise with RHEL, Ubuntu, Rocky, or similar enterprise Linux distributions.
· Strong troubleshooting skill across boot flow, system logs, networking stack, authentication, service lifecycle, and hardware-software interaction.
· Experience with GPU servers, out-of-band management, firmware coordination, and cluster node bring-up.
· Hands-on knowledge of Ansible, PXE/iPXE, Kickstart, cloud-init, image lifecycle management, and configuration enforcement.
· Strong shell scripting and Python-based automation capability.
· Working knowledge of storage and network dependencies affecting Linux host health.
· Ability to operate independently in ambiguous, high-severity production situations.
PREFERRED EXPERIENCE
· Exposure to Slurm, Kubernetes, container runtimes, or AI cluster schedulers.
· Familiarity with DCGM, Mellanox networking, and telemetry-driven health analysis.
Experience supporting validation labs or pre-production cluster certification
🔢 Crunching numbers...
Santa Clara, California
•
Today
WHAT THIS CANDIDATE WILL BE DOING Administer large-scale Linux environments supporting AI training, inference, and HPC workloads. Own deep troubleshooting of OS, kernel, boot, package, firmware, driver, filesystem, service, and resource-consumption issues across bare-metal server fleets. Diagnose failures across BIOS, BMC, PXE, DHCP, DNS, NFS, local disk, RAID,NVMe,systemd, and GPU driver stacks. Build andmaintaingolden images, provisioning pipelines, configuration baselines, and post-deployment
Easy Apply
Contract
50 - 60
Santa Clara, California
•
Today
Senior Linux Administrator AI & HPC InfrastructureLocation: Santa Clara, CA Onsite Job Type: Contract Experience: 7+ Years Domain: AI / HPC / Data Center Infrastructure Position OverviewWe are seeking a Senior Linux Administrator to support large-scale AI and High-Performance Computing (HPC) environments. The ideal candidate will be a hands-on Linux expert with strong experience troubleshooting production server fleets, GPU infrastructure, high-performance storage, and data center-connected co
Easy Apply
Contract
Depends on Experience
Sunnyvale, California
•
17d ago
Role: Hardware Systems Administrator (Linux) Location:Sunnyvale, CA - Onsite Experience: 5+ years Focus: Linux, Server Hardware & Infrastructure Key Requirements: Strong Linux administration experience (RHEL, Ubuntu, CentOS, Rocky Linux) Hands-on server hardware installation, configuration, upgrades, and troubleshooting Bash/Shell/Python scripting and automation Experience with storage, RAID, SAN/NAS, backups, and virtualization (VMware/KVM/Hyper-V) Strong networking knowledge: TCP/IP, DNS,
Easy Apply
Contract, Third Party
$70 - $75
Fremont, California
•
7d ago
Maxonic maintains a close and long-term relationship with our direct client. In support of their needs, we are looking for aSystems Admin (Linux Admin/SaaS and Windows Exp). Job Description: Job Title:Systems Admin (Linux Admin/SaaS and Windows Exp) Job Type:Fulltime Job Location:Fremont, CA Work Schedule:Onsite 5 days a week TheSystems Engineering team is responsible for architecting, building, automating, and managing our server Infrastructure at our campus in Fremont, CA and our public clou
Easy Apply
Full-time
Depends on Experience