Senior / Lead L2 OpenShift (OCP) Support Engineer
Iselin, NJ / NYC / Charlotte, NC
JD:
We are seeking a Senior / Lead L2 OpenShift (OCP) Support Engineer to drive the stability, performance, and operational excellence of our enterprise container platforms. In this hybrid role, you will act as the highest technical escalation point within the L2 tier, managing complex hybrid-cloud clusters spanning On-Premises, Microsoft Azure, and Google Cloud Platform (Google Cloud Platform).
This position requires extensive hands-on production support and Cloud Operations (Cloud Ops) experience in large-scale enterprise environments. As a lead, you will handle complex cluster and workload failures, oversee cloud infrastructure health, mentor junior engineers, and optimize operational efficiencies across diverse infrastructures.
Work Model & Location
- Type: Hybrid (Combination of remote and designated on-site office days).
- Location: Iselin, NJ or Charlotte, NC
Key Responsibilities
- Technical Escalation: Serve as the primary L2 escalation point for critical production incidents, driving rapid isolation and resolution under high-pressure scenarios.
- Hybrid-Cloud & Cloud Ops Administration: Monitor and maintain OpenShift clusters deployed across On-Premises bare-metal/VMware, Azure (ARO), and Google Cloud Platform, managing underlying cloud infrastructure lifecycles.
- Team Leadership & Mentorship: Guide, train, and mentor junior and mid-level L2 support engineers, ensuring high quality of service and technical growth across the team.
- Queue & SLA Management: Oversee the incident queue, assign tickets, and ensure the team meets strict Service Level Agreements (SLAs).
- Advanced Troubleshooting: Diagnose complex cluster-wide infrastructure failures, multi-cloud ingress/egress bottlenecks, security context constraints (SCCs), and hybrid persistent storage issues.
- Operational Efficiency: Collaborate on Cloud Ops initiatives including cloud cost optimization, resource utilization monitoring, IAM permission management, and infrastructure patching.
- Process Improvement: Lead the creation of standard operating procedures (SOPs), runbooks, and automated health-check scripts to improve team efficiency.
- Cross-Functional Collaboration: Partner closely with DevOps, Cloud Ops, Development, and L3 Platform Architecture teams to provide actionable operational feedback and ensure smooth application onboarding.
Required Qualifications & Experience (Must-Haves)
- Experience: Minimum of 8+ years in enterprise production support, cloud infrastructure engineering, or Cloud Ops, with at least 1–2 years in a senior or lead capacity.
- Hands-on Hybrid OCP: Proven, extensive hands-on experience administering, securing, and troubleshooting multi-cluster Red Hat OpenShift environments across On-Premises, Azure, and Google Cloud Platform.
- Cloud Operations Expertise: Strong operational knowledge of public cloud environments (Azure/Google Cloud Platform), including provisioning resources, managing cloud networking (VPCs, VNets, Peering), IAM roles, and cloud cost/governance tracking.
- Expert Container Orchestration: Deep knowledge of Kubernetes/OpenShift internals, API mechanics, operators, stateful applications, and cluster storage networking.
- Linux & Networking Expert: Advanced Linux administration skills (RHEL/CoreOS), with strong proficiency in diagnosing network layers, certificates, and OS-level resource constraints.
- Leadership Mindset: Strong ownership, exceptional communication skills, and the ability to coordinate incident bridge calls involving multiple stakeholders.
Preferred Qualifications
- Certification: Red Hat Certified Specialist in OpenShift Administration (EX280), Red Hat Certified Specialist in OpenShift Automation and DevOps (EX440), or CKA/CKAD is highly preferred.
- Cloud Certifications: Associate or Professional level certifications in Microsoft Azure (e.g., Azure Administrator) or Google Cloud Platform (e.g., Associate Cloud Engineer).
- Enterprise Observability: Experience designing or customizing monitoring dashboards using Prometheus, Grafana, ELK/Splunk, and setting up complex cloud alerting rules.
- Automation Engineering: Experience writing advanced Bash or Python scripts or using Ansible/Terraform to automate incident triage and platform verification.