Global Lead, Infrastructure Operations

Milpitas, CA, US • Posted 1 day ago • Updated 1 day ago
Contract Independent
Contract W2
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Microsoft Hyper-V
  • virtualization
  • Storage

Summary

Role: Global Lead, Infrastructure Operations

Location: Milpitas / Irvine, CA

Employment Type: Contract to hire


Scope: Global Infrastructure Operations

Role Summary

The Global Lead, Infrastructure Operations will own the global operating model, service health, and day-to-day coordination of enterprise infrastructure services across on-premises data centers and hybrid multicloud environments.

This is a working leadership role for an individual who can lead geographically distributed teams, establish operational discipline, and remain directly engaged in incident resolution, service improvement, and automation. The role requires broad infrastructure knowledge and the ability to coordinate technical specialists during complex, business-impacting events.

Working closely with the Chief Infrastructure Architect and domain leaders across compute, virtualization, storage, networking, cloud, user productivity, and cyber resilience, this leader will establish consistent global processes for monitoring, incident response, change execution, escalation, and service reporting.

Domain leaders retain ownership of their platforms and engineering standards. The Global Operations Lead owns coordinated service delivery, operational visibility, and cross-domain response.

Key Responsibilities

Global Operations Leadership

  • Define and lead a global infrastructure operating model, including regional coverage, follow-the-sun handoffs, on-call responsibilities, and escalation procedures.
  • Lead operations staff and coordinate regional engineering teams, managed service providers, and technology partners.
  • Establish service ownership, support boundaries, operational procedures, and measurable service expectations.
  • Maintain a consolidated view of infrastructure service health, business impact, operational risks, and outstanding remediation.
  • Own operations staffing plans, service-provider performance, and operational tooling requirements.
  • Build a culture of accountability, clear communication, and continuous improvement.

Incident, Problem, and Change Management

  • Lead infrastructure incident response, bringing together the appropriate technical teams and maintaining clear ownership through resolution.
  • Serve as the infrastructure incident commander during major service disruptions, coordinating technical response, stakeholder updates, and recovery verification.
  • Partner with cybersecurity incident response during cyber events and with resilience leaders during recovery activities.
  • Establish incident severity definitions, escalation thresholds, communication procedures, and response expectations.
  • Lead post-incident reviews and ensure corrective actions address recurring issues and underlying causes.
  • Coordinate the global change calendar, maintenance windows, dependency reviews, and operational readiness for significant changes.
  • Ensure changes include implementation plans, rollback procedures, appropriate approvals, and post-change validation.

Observability and Service Reliability

  • Establish monitoring and observability requirements across infrastructure platforms and critical business services.
  • Partner with domain and application teams to connect infrastructure events to service dependencies and business impact.
  • Improve alert quality through correlation, deduplication, actionable thresholds, and clear routing.
  • Define and report service-level indicators, service-level objectives, availability, response times, and restoration performance.
  • Drive proactive identification of capacity constraints, performance degradation, configuration drift, and recurring failure patterns.
  • Advance operations from reactive monitoring toward predictive and preventive practices.

Operational Automation and AIOps

  • Develop an automation roadmap for incident enrichment, diagnostics, service requests, routine remediation, and operational reporting.
  • Partner with engineering teams to implement version-controlled runbooks and workflows using scripts, APIs, and automation platforms.
  • Evaluate and implement AIOps and agentic automation for event analysis, troubleshooting assistance, and controlled remediation.
  • Establish least-privilege access, approval thresholds, audit trails, testing, and rollback requirements for automated actions.
  • Integrate observability, IT service management, configuration data, and automation to improve response and reduce manual work.
  • Measure automation effectiveness through reduced operational effort, faster restoration, and fewer repeated incidents.

Hybrid Infrastructure Operations

  • Coordinate operational support across Dell compute and storage, Microsoft Hyper-V, Nutanix AHV, and hybrid Nutanix environments, including NC2.
  • Establish consistent support practices across on-premises infrastructure and AWS, Microsoft Azure, and Google Cloud Platform.
  • Partner with domain leaders to track patch compliance, backup health, replication status, capacity, platform lifecycle, and vulnerability remediation.
  • Ensure service dependencies, support contacts, escalation procedures, and recovery runbooks remain accurate.
  • Coordinate operational participation in disaster recovery, cyber recovery, failover, and failback exercises.
  • Validate service stability following upgrades, migrations, major changes, and recovery events.

Service Transition and Governance

  • Define operational acceptance criteria for new platforms and services, including monitoring, documentation, training, support coverage, and recovery procedures.
  • Ensure project teams complete effective handoffs to operations before production acceptance.
  • Maintain a service catalog and support accurate asset, configuration, and service dependency records.
  • Establish dashboards and service reviews that connect operational performance to business priorities.
  • Manage service-provider commitments, escalation performance, and improvement plans.
  • Communicate operational risks, resource needs, and improvement priorities to infrastructure leadership.

Required Qualifications

  • 10+ years of progressive experience in enterprise infrastructure operations, engineering, or service delivery, including leadership of global or geographically distributed teams.
  • Demonstrated experience owning infrastructure service performance and coordinating operations across multiple technical domains.
  • Proven ability to lead major incident response in complex enterprise environments.
  • Strong practical knowledge of incident, problem, change, configuration, and service-level management.
  • Broad technical understanding of compute, virtualization, storage, networking, cloud, backup, and recovery.
  • Experience supporting environments using Microsoft Hyper-V, Nutanix, Dell compute, and Dell enterprise storage.
  • Experience operating hybrid environments spanning on-premises infrastructure and public cloud, with familiarity with AWS, Azure, and Google Cloud Platform.
  • Hands-on ability to interpret monitoring data, logs, infrastructure health indicators, and service dependencies to guide troubleshooting.
  • Experience with enterprise observability and IT service management platforms, including dashboarding, alert management, escalation, and reporting.
  • Experience implementing automation through scripting, APIs, runbooks, or orchestration platforms.
  • Experience establishing global support coverage, regional handoffs, on-call models, and vendor escalation processes.
  • Strong executive communication skills, including the ability to explain business impact, response progress, and operational risk during major incidents.
  • Demonstrated ability to lead through influence while maintaining clear accountability across technical teams.

Preferred Qualifications

  • Experience in semiconductor, manufacturing, or other global environments requiring continuous infrastructure availability.
  • Familiarity with Nutanix Prism and NC2, including operational dependencies across hybrid environments.
  • Experience with tools such as ServiceNow, Dynatrace, Splunk ITSI, Grafana, SCOM, NetBrain, or comparable platforms.
  • Experience implementing AIOps, event correlation, predictive operations, or agentic automation.
  • Familiarity with PowerShell, Python, Ansible, Terraform, and automated operational workflows.
  • Familiarity with NetBackup, Rubrik, or Cohesity, including backup health monitoring and recovery escalation.
  • Experience applying site reliability engineering practices, including service-level objectives, reduction of repetitive manual work, and blameless incident reviews.
  • Experience transitioning infrastructure services from project delivery into steady-state operations.
  • ITIL certification or relevant cloud, infrastructure, or service management credentials.

Measures of Success

  • Improved service availability and faster detection, response, and restoration.
  • Consistent global coverage, effective regional handoffs, and clear service ownership.
  • Fewer recurring incidents and timely completion of corrective actions.
  • Improved change success rates and fewer change-related disruptions.
  • Reduced alert noise and increased visibility into business-service health.
  • Greater automation of routine work and repeatable operational procedures.
  • Reliable operational handoffs for new platforms and modernization initiatives.

Clear, actionable reporting that enables leadership to prioritize service improvements.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91126849
  • Position Id: 9097000
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Milpitas, California

•

Yesterday

Easy Apply

Contract

Depends on Experience

Cupertino, California

•

3d ago

Full-time

Cupertino, California

•

3d ago

Full-time

Menlo Park, California

•

Today

Easy Apply

Full-time

GBP - GBP

Search all similar jobs