Sr. Staff Site Reliability/SRE _Remote_$70 per hour on W2_10+ yrs exp


Xoriant Corporation
Dice Job Match Score™
🧠 Analyzing your skills...
Job Details
Skills
- SRE/Site Reliability
- Production on-call
- Kubernetes
- Docker
- AWS/GCP/Azure
- Terraform/OpenTofu
- Prometheus/Grafana /Observability /Monitoring
- Artificial intelligence/AI/Machine Learning/ML
- PYTHON OR BASH OR GO
- CICD
Summary
Xoriant is an equal opportunity employer. No person shall be excluded from consideration for employment because of race, ethnicity, religion, caste, gender, gender identity, sexual orientation, marital status, national origin, age, disability or veteran status.
TITLE:- Sr. Staff SRE (AI/ML)
LOCATION Remote
DURATION 12+ Months (May get extend)
MODE OF INTERVIEW Zoom/Webex
RATE $70 per hour on W2
JOB DESCRIPTION
- Production on-call experience in a real rotation, with incident command and blameless postmortem practice.
Production Kubernetes and container experience (Docker), with cloud-native infrastructure patterns.
Hands-on production ownership on at least one major cloud (AWS, Google Cloud Platform, or Azure).
Terraform or OpenTofu proficiency.
Observability depth with Prometheus, Grafana, or equivalent for metrics, logging, and alerting, including dashboard and alert design.
Strong automation skills in Python, Bash, or Go.
Networking fundamentals: VPCs, load balancers, DNS, firewalls, cross-cloud connectivity.
CI/CD experience with GitHub Actions, GitLab CI, Jenkins, or ArgoCD.
Proven ability to troubleshoot complex distributed systems, largely self-directed.
Preferred Qualifications GPU infrastructure and AI/ML workloads: Ray, Kubeflow, ML flow, or similar.
NVIDIA GPU orchestration: A100/H100 configuration, driver and CUDA runtime management.
Distributed training networking: RDMA, InfiniBand, EFA, NCCL.
Distributed tracing and Open Telemetry instrumentation across services.
Progressive delivery: canary and blue/green rollouts with automated rollback.
Chaos or fault-injection testing, game days, and disaster-recovery drills.
Multi-cloud networking, unified storage abstractions, and disaster recovery.
FinOps and cost optimization: Spot, Reserved Instances, Savings Plans.
Establishing an SRE function where one did not previously exist.
///****Any query can call on four zero eight five five zero one two eight seven*******///////////////
- Dice Id: xorca001
- Position Id: 9051598
- Posted 6 hours ago
Company Info
Xoriant is a Sunnyvale, CA headquartered digital engineering firm with offices in the USA, Europe, and Asia. From Tech Startups to Fortune 100 Enterprises, we enable innovation, accelerate time to market, and ensure client competitiveness across industries. Across all our focus areas – platform engineering, cloud, data & and AI, and Security – every solution we develop benefits from our product engineering DNA and culture of innovation. It also includes successful methodologies, framework components, and accelerators for rapidly solving critical client challenges. For 30 years and counting, we have taken great pride in the longlasting, deep relationships we have with our clients.
For further information about Xoriant, please visit our website

Similar Jobs
It looks like there aren't any Similar Jobs for this job yet.
Search all similar jobs