Senior/Staff SRE for AI/ML Platform Infrastructure

San Jose, CA, US • Posted 12 hours ago • Updated 12 hours ago
Contract Independent
Contract Corp To Corp
Contract W2
12 Months
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🎯 Assessing qualifications...

Job Details

Skills

  • Amazon Web Services
  • Artificial Intelligence
  • CHAOS
  • CUDA
  • Bash
  • Cloud Computing
  • Computer Networking
  • Continuous Delivery
  • Continuous Integration
  • DNS
  • Dashboard
  • GPU
  • GitHub
  • Disaster Recovery
  • Dragon NaturallySpeaking
  • Firewall
  • Docker
  • Good Clinical Practice
  • Google Cloud Platform
  • Grafana
  • GitLab
  • InfiniBand
  • Instrumentation
  • Jenkins
  • Kubernetes
  • Optimization
  • Orchestration
  • Remote Direct Memory Access
  • Machine Learning (ML)
  • Microsoft Azure
  • Management
  • Storage
  • Terraform
  • Testing
  • Python

Summary

Minimum Qualifications
 Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
 Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
 Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
 Terraform or OpenTofu proficiency.
 Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
 Strong automation skills in Python, Bash, or Go.
 Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
 CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
 Proven ability to troubleshoot complex distributed systems, largely self-directed.


Preferred Qualifications
 GPU infrastructure and AI/ML workloads: Ray, Kubeflow, MLflow, or similar.
 NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
 Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
 Distributed tracing and OpenTelemetry instrumentation across services.
 Progressive delivery: canary and blue/green rollouts with automated rollback.
 Chaos or fault-injection testing, game days, and disaster-recovery drills.
 Multi-cloud networking, unified storage abstractions, and disaster recovery.
 FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
 Establishing an SRE function where one did not previously exist.

 

Thanks & Regards,

Narendra Kunware

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: xorca001
  • Position Id: 9052791
  • Posted 12 hours ago

Company Info

About Xoriant Corporation

Xoriant is a Sunnyvale, CA headquartered digital engineering firm with offices in the USA, Europe, and Asia. From Tech Startups to Fortune 100 Enterprises, we enable innovation, accelerate time to market, and ensure client competitiveness across industries. Across all our focus areas – platform engineering, cloud, data & and AI, and Security – every solution we develop benefits from our product engineering DNA and culture of innovation. It also includes successful methodologies, framework components, and accelerators for rapidly solving critical client challenges. For 30 years and counting, we have taken great pride in the longlasting, deep relationships we have with our clients.

For further information about Xoriant, please visit our website

About_Company_One
Contact the job poster
Narendra Kunware

Narendra Kunware

Recruiter @ Xoriant Corporation
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Yesterday

Easy Apply

Contract

$60 - $70

San Jose, California

12d ago

Easy Apply

Contract, Third Party

75 - 80

Search all similar jobs