Job Title: System Test Engineer
Preferred Location: Highly desired in Santa Clara , CA , or Costa Rica (Lower Cost Center Preferred)
Duration: Long term Contract
About the role
Storage platform for AI factories a fleet of metadata nodes (MDN) and commodity data nodes (DN) connected over pNFS, RDMA, and gRPC control channels, running a EverPure Purity//DN OS on off-the-shelf hardware. That combination heterogeneous hardware, a custom OS image, a control/data-plane split, and deployment into performance-sensitive HPC/AI and NeoCloud environments means system-level validation is as critical to //EXA's competitiveness as the performance engineering itself: a fleet that can't be reliably configured, upgraded, and diagnosed in the field undermines every throughput and latency claim.. This role owns that validation layer: proving the platform holds up as real hardware, firmware, networking, and operating system state, not just as an architecture diagram.
What you'll do
- System Testing for the FlashBlade//EXA product line.
- Validate software configuration end-to-end. Test configuration management across MDN and DN roles geometry construction/distribution, export and pack-group configuration, node onboarding ensuring config changes propagate correctly and safely across the fleet.
- Test networking at a working-knowledge depth. Validate the system's network-dependent behavior: RDMA/data-path connectivity between clients and DNs, control-channel (gRPC/TCP) reachability, VIP/failover behavior, and network-induced fault conditions (link flaps, partition, degraded fabric) enough networking depth to design and troubleshoot these test scenarios independently.
- Filing and owning related bugs (documenting, debugging, follow-up with Dev teams, reproduce, etc)
- Own and drive features to completion (review requirements, design docs, customer use cases to build test strategies, test cases, and end-to-end workflows in the System Test ecosystem for new and existing features over multiple releases)
- Build up and scale-out of complex system testbeds and test ecosystems using virtualized environments and various test tools.
- Build and maintain CI/CD test infrastructure. Own Jenkins pipelines and Docker-based test environments for system and platform test suites reproducible test builds, scheduled regression runs, and fast feedback for engineering.
- Automate via scripting. Write and maintain Python and Shell tooling for test orchestration, environment setup/teardown, log collection, and result validation across MDN/DN clusters.
- Own foundation / platform testing. Validate the base platform underneath the data-plane features DNOS image provisioning, node bring-up, hardware/OS compatibility, platform-level health and telemetry the layer everything else in //EXA depends on.
- Test firmware upgrade and maintenance workflows. Validate FW upgrade paths on data-node hardware (including Viking dual-node/HA configurations) for correctness, rollback safety, and no-downtime operation upgrades must not compromise the HA guarantees the platform advertises.
- Work on Customer Escalations with cross-functional teams to help identify root causes and provide assistance with reproductions.
- Troubleshoot at the linux kernel level. Diagnose failures directly on DNOS/Linux hosts kernel logs, systemd/service state, filesystem and storage stack issues, network stack behavior and turn root causes into actionable bug reports or regression tests.
What makes you competitive for this role
- Experience in System Testing with the ability to create end-to-end test ecosystems from the ground up and maintain it over time. Understand the nature of a workflow vs a functional test
- Knowledge of flash storage, file systems/protocols: Flash media, Read/Write IOPS, IO datapath, NFS, SMB, S3, and networking layers (full-stack QA)
- Solid software configuration management experience comfortable validating config propagation and drift in a distributed, multi-node system.
- Mid-level networking knowledge: TCP/IP fundamentals, RDMA/RoCE or InfiniBand exposure a plus, comfortable diagnosing connectivity and fabric-level issues.
- Experience with switching/fabric configuration (Nvidia/Mellanox, Cisco, Arista) in the areas of BGP, ECMP, VLAN, RoCE
- Hands-on DevOps skills Jenkins pipeline authoring/maintenance and Docker-based test environments. IaC orchestration (Ansible, Terraform)
- Strong scripting ability in Python and Shell for test automation and tooling.
- Experience with foundation/platform-level testing (OS image, hardware bring-up, base system health, component level HA, fault injection, stress testing) rather than only feature-level testing.
- Direct experience testing firmware upgrade and maintenance procedures, ideally on multi-node or HA hardware.
- Strong Linux fundamentals and troubleshooting instincts able to root-cause from logs and system state rather than guesswork.
- Designing System Test Strategy (System test Plans, Topologies, Coverage design)
- Performance / Load / Workload Simulation over NFSv4.1, NFSv3 and S3 protocols
- Experience in building out and working with Virtualized environments is a plus (VMware, KVM, OpenStack)