ENGAGEMENT SUMMARY
The Candidate will provide senior network engineering services for Ethernet-based AI data center fabrics
supporting GPU compute clusters, storage connectivity, and management plane services. The role requires
strong operational judgment, hands-on troubleshooting, and the ability to resolve failures across spine-leaf
architectures under production pressure.
WHAT THIS CANDIDATE WILL BE DOING
• Deploy, validate, and support Ethernet fabrics used for AI and HPC cluster environments.
• Troubleshoot L1 through L4 issues involving optics, transceivers, cabling, link bring-up, VLANs, MLAG,
ECMP, BGP, underlay/overlay reachability, congestion, and packet loss.
• Validate network readiness for distributed training and large east-west traffic patterns.
• Diagnose performance issues related to buffer pressure, microbursts, PFC behavior, QoS policy, MTU
mismatch, routing instability, and oversubscription.
• Work closely with Linux, storage, and cluster deployment teams to isolate host-versus-network fault
domains.
• Review and execute change plans for switch provisioning, firmware upgrades, topology expansion, and
maintenance events.
• Capture packet-level and counter-based evidence to drive root cause analysis.
• Build operational standards for cable mapping, port policy consistency, and fabric health validation.
WHAT WE NEE D TO SEE
• 7+ years in large-scale data center networking, including high-bandwidth Ethernet fabrics.
• Strong experience with spine-leaf design, routing, switching, and production troubleshooting.
• Hands-on skill with BGP, EVPN/VXLAN, MLAG, ECMP, QoS, PFC, RoCE considerations, and telemetry
interpretation.
• Experience validating optics, breakout configurations, cable plant integrity, and port-level consistency.
• Proven ability to troubleshoot distributed application impact caused by network behavior.
• Comfort using switch CLI, automation tooling, and packet/counter analysis workflows.
• Strong documentation habits for topology, incident timelines, and remediation plans.
PREFERRED EXPERIENCE
• Direct experience with AI fabrics carrying large-scale GPU collective traffic.
• Familiarity with SONiC, Cumulus Linux, or vendor NOS platforms used in AI data centers.
• Experience with streaming telemetry, PrometheGrafana, and network SRE operating models