W2 -Senior Site Reliability Engineer :- Santa Clara, CA(Onsite)

Santa Clara, CA, US • Posted 3 days ago • Updated 3 days ago
Contract Corp To Corp
Contract W2
12 Months
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • SRE
  • AI
  • Linux
  • Kubernetes

Summary

Role :- Senior Site Reliability Engineer

Location :- Santa Clara, CA(Onsite)


P O S I T I O N D E S C R I P T I O N

ENGAGEMENT SUMMARY


The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on

reliability engineering, incident response, service health, and operational automation. This role is best suited to

a senior hands-on engineer who can improve availability while remaining effective in detailed production

troubleshooting.

WHAT THIS CANDIDATE WILL BE DOING

• Operate and improve reliability of AI platform services, cluster dependencies, and shared infrastructure

components.

• Lead or support incident triage for service degradation involving Kubernetes, Linux hosts, storage, network,

scheduling, job orchestration, or dependency failures.

• Define and refine SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident actions.

• Analyze recurring failure patterns and convert manual operations into automation and preventive controls.

• Build observability across system, service, workload, and dependency layers using metrics, logs, traces, and

event correlation.

• Troubleshoot performance and availability issues affecting training jobs, inference services, internal

platforms, and support tooling.

• Partner with infrastructure and validation teams to improve production readiness and change safety.

• Drive operational reviews, readiness criteria, and resilience testing.

WHAT WE NEE D TO SEE

• 7+ years in SRE, production operations, or reliability-focused infrastructure engineering.

• Strong hands-on troubleshooting across Linux, Kubernetes, networking, and distributed systems.

• Experience building observability, alerting, and response workflows in complex production environments.

• Ability to balance urgent operational response with medium-term reliability engineering improvements.

• Strong scripting and automation skills, with experience reducing toil through tooling.

• Experience participating in incident management, root cause analysis, and post-incident follow-through.

• Strong communication skill with the ability to summarize technical issues clearly for cross-functional teams.

PREFERRED EXPERIENCE

• Experience in AI platforms, ML infrastructure, or large-scale HPC-like service environments.

• Familiarity with Prometheus, Grafana, ELK/OpenSearch, Loki, PagerDuty, and incident tooling.

• Experience defining error budgets and applying SRE practices in environments with heavy batch and service

traffic.


Thanks and Regards,
Akash

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91112163
  • Position Id: 8704-32090-1787934059
  • Posted 3 days ago

Company Info

About Noblesoft Technologies Inc.

Noblesoft's Executive Leadership is driven by a profound commitment to client success and employee growth. The team sets the strategic direction by focusing on mastery within core Enterprise Application practices—Salesforce, SAP, and Workday, among others. By prioritizing innovative solution delivery and fostering a culture where every team member is empowered, the leadership ensures Noblesoft remains a trusted partner in digital transformation across its specialized industries.

The leadership team's deep expertise in enterprise technology enables Noblesoft to navigate complex implementation challenges and deliver high-value, scalable solutions. They are instrumental in continuously evolving the company's offerings, ensuring services remain aligned with the latest advancements in CRM, ERP, and HCM platforms. This proactive approach to market trends ensures that Noblesoft consultants are always working with best-in-class methodologies to drive client modernization initiatives.

Beyond technical strategy, the core philosophy of the leadership is centered on people. They maintain an employee-centric focus, viewing the talented consultant base as the primary asset for ensuring project excellence and cultivating long-term client relationships. This dedication to internal growth includes sponsoring continuous professional development, specialized certifications, and mentorship programs across all major practices.

Ultimately, the vision of the Executive Leadership is to position Noblesoft not just as a technology implementer, but as a strategic advisory partner for global enterprises. Their guidance ensures every solution deployed—whether a Salesforce customization, a SAP S/4HANA migration, or a Workday deployment—is rooted in a clear understanding of the client's business objectives, securing the company's reputation for great products, sustainable solutions, and enduring client partnerships.

Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs