REQUIRED SKILLSET
Candidate must demonstrate experience building, designing, or operating distributed storage systems or components. Simply using a distributed database or storage system as an application backend does not satisfy this requirement. Candidate should be able to explain concepts such as data consistency, replication behavior, and whether data written to one node can be immediately observed from another node, supported by examples from prior projects. Preferably located in the Bay Area or San Jose, but not required.
KEY RESPONSIBILITIES
1. Design, develop, maintain and provide on-call support for high-performance, scalable, and reliable backend services that interact with various storage systems
2. Troubleshoot production issues, perform debugging, and provide operational support
REQUIRED SKILLS
1. Coding:
-
- Strong hands-on experience in Rust with proven backend development expertise.
- Working knowledge of Golang, Java, and Python, demonstrating the ability to quickly onboard to new frameworks and begin developing non-trivial features in these languages within two days.
- Knowledge of C/C++, to understand complex systems implemented in these languages.
2. Design:
-
- Strong foundation in Data Structures and Algorithms (DSA), including trees, graphs, hash tables, heaps, dynamic programming, and algorithmic problem solving.
- Solid experience in system design, especially including designing scalable, distributed, and high-availability backend systems.
3. Scheduling:
-
- Candidate must be flexible in supporting collaboration across global teams, including occasional early morning or late evening meetings to overlap with teams in the America/Los Angeles and Asia/Shanghai timezones.
- Core collaboration hours are expected to maximize availability between 9:00 AM and 6:00 PM in either timezone, adjusted for Daylight Saving Time as applicable.
REQUIRED TECHNICAL EXPERIENCE:
Candidate must demonstrate hands-on experience across all categories listed below, with at least one relevant technology in each category supported by prior project or assignment experience. A single project or assignment may be used to demonstrate expertise across multiple categories.
Concurrent Programming: Experience with asynchronous programming, multi-threading, concurrency models, and related design patterns.
Remote Procedure Call (RPC): Experience designing or implementing RPC-based services using technologies such as gRPC or Apache Thrift. (RESTful HTTP APIs alone do not satisfy this requirement.)
Cloud Native Environment: Experience with cloud-native service architectures, including:
- Service discovery technologies such as Consul or ZooKeeper (excluding DNS-based discovery only)
- Service mesh technologies such as Envoy or Istio
- IPv6 networking concepts and implementation
Distributed Storage Systems: Strong understanding of distributed storage concepts, including:
- Data consistency models and trade-offs
- Data partitioning/sharding strategies, including routing and resharding
- Replication mechanisms and fault tolerance
- Distributed system behavior and scalability considerations