Xoriant is an equal opportunity employer. No person shall be excluded from consideration for employment because of race, ethnicity, religion, caste, gender, gender identity, sexual orientation, marital status, national origin, age, disability or veteran status.
[**** NOTE - NO RELOCATION CANDIDATE , ONLY LOCAL TO SanJose, CA , Who is ready to attend face to face interview*********]
TITLE:- Sr. Staff SRE (AI/ML)
LOCATION - SanJose, CA
DURATION 12+ Months (May get extend)
MODE OF INTERVIEW- Face to Face must
RATE - $65 per hour on C2C
JOB DESCRIPTION
- Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
Terraform or OpenTofu proficiency.
Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
Strong automation skills in Python, Bash, or Go.
Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
Proven ability to troubleshoot complex distributed systems, largely self-directed.
Preferred Qualifications GPU infrastructure and AI/ML workloads: Ray, Kubeflow, ML flow, or similar.
NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
Distributed tracing and Open Telemetry instrumentation across services.
Progressive delivery: canary and blue/green rollouts with automated rollback.
Chaos or fault-injection testing, game days, and disaster-recovery drills.
Multi-cloud networking, unified storage abstractions, and disaster recovery.
FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
Establishing an SRE function where one did not previously exist.