Role: Manager, Data Center & Infrastructure Operations Lead
Location: Milpitas / Irvine, CA Hybrid
Employment Type: Contract to hire
Position Overview
We are seeking an experienced Manager, Data Center & Infrastructure Operations Lead to oversee the day-to-day operations, availability, performance, and reliability of enterprise data center and infrastructure environments.
This role will lead infrastructure operations teams responsible for data center systems, servers, storage, virtualization, network infrastructure, facilities coordination, monitoring, incident management, and operational support. The ideal candidate is a strong technical leader who can manage both people and complex infrastructure environments while driving operational excellence, resiliency, and continuous improvement.
The position will work closely with Infrastructure Engineering, Network, Storage, Cloud, Security, Applications, Facilities, and Service Management teams to ensure highly available and reliable IT services.
Key Responsibilities
- Lead day-to-day operations of enterprise data center and infrastructure environments.
- Manage infrastructure operations teams responsible for servers, storage, virtualization, data center systems, and related technologies.
- Ensure availability, performance, capacity, and reliability of critical infrastructure.
- Establish and maintain operational procedures, standards, runbooks, and documentation.
- Monitor infrastructure health and proactively identify performance, capacity, and availability issues.
- Lead response and escalation for infrastructure incidents and service disruptions.
- Coordinate major incident response and drive issues through resolution.
- Conduct root-cause analysis and develop corrective and preventative actions.
- Manage infrastructure maintenance, upgrades, patching, hardware refreshes, and lifecycle activities.
- Coordinate planned maintenance windows and minimize business impact.
- Partner with Network, Storage, Cloud, Security, and Application teams on infrastructure initiatives.
- Work closely with Data Center Facilities teams and vendors to coordinate power, cooling, cabling, physical security, and environmental requirements.
- Maintain accurate infrastructure inventories, configurations, diagrams, and operational documentation.
- Develop and manage infrastructure capacity plans and technology refresh roadmaps.
- Establish and monitor operational KPIs and service-level objectives.
- Ensure infrastructure environments meet availability, resiliency, security, and compliance requirements.
- Support disaster recovery and business continuity initiatives, including infrastructure recovery testing.
- Identify opportunities to automate repetitive operational processes and improve infrastructure efficiency.
- Manage relationships with hardware, software, data center, and managed-service vendors.
- Assist with budgeting, forecasting, procurement, and infrastructure lifecycle planning.
Data Center Operations
- Oversee day-to-day operation of enterprise data center environments.
- Coordinate data center equipment installation, decommissioning, relocation, and refresh activities.
- Ensure proper rack, power, cooling, cabling, and physical infrastructure standards.
- Partner with facilities teams on power, HVAC, UPS, generators, environmental monitoring, and physical security.
- Maintain data center standards, procedures, documentation, and emergency response processes.
- Coordinate data center moves, adds, changes, and infrastructure deployments.
- Monitor data center capacity and identify potential power, cooling, space, and infrastructure constraints.
Infrastructure Operations
- Oversee physical and virtual server infrastructure.
- Manage enterprise virtualization environments and infrastructure platforms.
- Coordinate storage and backup operations with appropriate engineering teams.
- Ensure infrastructure monitoring and alerting systems are operating effectively.
- Manage operating system and infrastructure patching processes.
- Coordinate hardware replacement and technology lifecycle management.
- Maintain infrastructure availability and performance standards.
- Ensure operational readiness for new infrastructure deployments.
Incident & Problem Management
- Serve as an escalation point for complex infrastructure incidents.
- Lead technical teams during major incidents and service-impacting events.
- Coordinate communications with IT leadership and business stakeholders.
- Ensure incidents are properly documented and resolved within established SLAs.
- Lead post-incident reviews and root-cause analysis.
- Identify recurring issues and implement permanent corrective actions.
- Partner with Service Management teams to improve operational processes.
Leadership & Team Management
- Lead, mentor, and develop infrastructure operations engineers and technical staff.
- Establish clear responsibilities, performance expectations, and operational standards.
- Manage staffing, schedules, on-call rotations, and operational coverage.
- Promote a culture of accountability, collaboration, automation, and continuous improvement.
- Provide technical guidance and career development for team members.
- Manage third-party contractors and managed-service resources as needed.
Required Qualifications
- 7+ years of experience in IT infrastructure, data center, systems, or infrastructure operations.
- 3+ years of experience managing infrastructure or data center operations teams.
- Strong understanding of enterprise data center environments and infrastructure operations.
- Experience with physical and virtual server environments.
- Experience with enterprise storage, backup, and networking technologies.
- Strong understanding of virtualization technologies, preferably VMware.
- Experience with infrastructure monitoring, alerting, incident management, and operational processes.
- Experience managing infrastructure hardware refreshes and technology lifecycle programs.
- Experience with data center facilities, including power, cooling, UPS, generators, and environmental monitoring.
- Strong understanding of ITIL-based operational processes.
- Experience managing vendors and third-party service providers.
- Strong troubleshooting, analytical, and problem-solving skills.
- Excellent communication and team leadership abilities.
Preferred Qualifications
- Experience with Dell server and infrastructure technologies.
- Experience with Dell storage platforms.
- Experience with Rubrik or other enterprise backup technologies.
- Experience with VMware vSphere and enterprise virtualization.
- Experience with Windows and Linux server environments.
- Experience with enterprise monitoring tools.
- Experience with data center migration or consolidation projects.
- Experience supporting hybrid cloud infrastructure.
- Familiarity with AWS, Azure, or other cloud platforms.
- Experience with automation and scripting, including PowerShell or Python.
- ITIL certification or equivalent experience.
- Bachelor s degree in information technology, Computer Science, Engineering, or related field preferred.
Key Success Measures
- High availability and reliability of critical infrastructure environments.
- Infrastructure incidents are resolved quickly and effectively.
- Data center operations meet established availability, security, and safety standards.
- Infrastructure capacity and lifecycle requirements are proactively managed.
- Maintenance and technology refresh activities are completed with minimal business disruption.
- Infrastructure documentation and operational procedures remain accurate and current.
- Operational processes become increasingly automated and efficient.
- Infrastructure teams maintain strong collaboration with Engineering, Security, Network, Storage, Cloud, and Application teams.