
Stryde Consulting Services LLC
Santa Clara, California • Yesterday
Easy Apply
Contract
Depends on Experience
108 results (3 new)

Stryde Consulting Services LLC
Santa Clara, California • Yesterday
Easy Apply
Contract
Depends on Experience


Xoriant Corporation
San Jose, California • 16d ago
Easy Apply
Contract, Third Party
Depends on Experience




















Stryde Consulting Services LLC
Santa Clara, California • Today
Easy Apply
Contract
Depends on Experience




Location: Santa Clara, CA – Onsite
Job Type: W2 Contract
Experience: 7+ Years
Domain: AI Infrastructure / Kubernetes / Cloud Platform
We are seeking a Senior Kubernetes Platform Engineer to support AI infrastructure environments focused on model development, distributed training, inference services, and shared platform operations.
The ideal candidate will have deep hands-on Kubernetes administration and troubleshooting experience, strong Linux fundamentals, and experience supporting GPU-enabled workloads in bare-metal and data center environments.
This is a hands-on platform engineering role requiring the ability to troubleshoot complex issues across Kubernetes, compute, networking, storage, containers, and GPU infrastructure.
Build, administer, maintain, and troubleshoot Kubernetes platforms supporting AI and data-intensive workloads.
Troubleshoot Kubernetes control plane components, CNI, CSI, ingress, service discovery, scheduling, node lifecycle, container runtimes, and resource isolation.
Support GPU-enabled Kubernetes environments, including device plugins, GPU drivers, node health, resource allocation, and workload placement.
Troubleshoot cluster and workload issues involving storage performance, networking, DNS, network policies, image pulls, autoscaling, pod evictions, and degraded nodes.
Perform root-cause analysis across nodes, pods, networking, storage, GPU infrastructure, and Kubernetes control plane components.
Improve platform reliability through automation, standardized configurations, upgrade planning, and cluster validation.
Manage Kubernetes cluster lifecycle activities, including upgrades, configuration changes, health checks, and validation.
Partner with Linux, network, storage, validation, and SRE teams to resolve complex cross-layer infrastructure issues.
Develop reusable operational runbooks, dashboards, monitoring, and health checks for day-2 operations.
Contribute to platform hardening, tenant readiness, reliability, and service-level objectives.
Automate platform administration and operational tasks using Python, Bash, or Go.
7+ years of infrastructure/platform engineering experience with strong hands-on Kubernetes administration.
Deep understanding of Kubernetes architecture and internals.
Strong experience troubleshooting Kubernetes clusters and production workloads.
Hands-on experience with:
Kubernetes
Container runtimes
Helm
CNI and CSI
Ingress
Kubernetes scheduling
Node lifecycle management
Cluster upgrades and lifecycle management
Experience with GPU workloads on Kubernetes in production, validation, lab, or AI infrastructure environments.
Strong Linux administration and troubleshooting skills.
Understanding of data center networking and storage dependencies.
Ability to troubleshoot issues from initial symptoms through root cause across multiple infrastructure layers.
Experience with declarative infrastructure, GitOps, or configuration-driven operations.
Strong scripting and automation skills using Python, Bash, or Go.
Strong communication and collaboration skills with infrastructure and engineering teams.
Ability to work independently in high-severity and ambiguous production situations.
Kubeflow
Argo / Argo CD
Prometheus
Grafana
Loki
Service mesh technologies
Bare-metal Kubernetes
High-performance storage
GPU cluster infrastructure
NVIDIA GPU ecosystem
AI/ML platform infrastructure
Distributed training environments
Regulated or high-change-control production environments
The ideal candidate is a hands-on Kubernetes platform engineer who can operate beyond standard cluster administration and troubleshoot complex AI infrastructure issues involving Kubernetes, Linux, GPUs, networking, storage, and compute.
Candidates should be comfortable working directly with bare-metal infrastructure and data center environments and collaborating across platform, Linux, networking, storage, validation, and SRE teams.
Santa Clara, CA – 100% Onsite
W2 engagement
Candidates must be comfortable supporting hands-on AI infrastructure and Kubernetes platform operations.
Stryde Consulting, an emerging leader in the HR Services domain, is promoted by Young professionals with years of industry experience with some of the top organizations in India. Incorporated in Hyderabad in the year 2005, Stryde aims to grow fast to become one of the top consulting firms in India, dealing with only the reputed and professional clients across industries, in India, Middle East and the US.

🔢 Crunching numbers...
Santa Clara, California
•
Today
Senior Linux Administrator AI & HPC InfrastructureLocation: Santa Clara, CA Onsite Job Type: Contract Experience: 7+ Years Domain: AI / HPC / Data Center Infrastructure Position OverviewWe are seeking a Senior Linux Administrator to support large-scale AI and High-Performance Computing (HPC) environments. The ideal candidate will be a hands-on Linux expert with strong experience troubleshooting production server fleets, GPU infrastructure, high-performance storage, and data center-connected co
Easy Apply
Contract
Depends on Experience