![]()
Software Engineer, ML Data Infrastructure
W2 Contract
Pay Rate: $60 - $70 per hour
Location: Cupertino, CA - Remote Role
Job Summary:
We are hiring a Software Engineer to build and run the data infrastructure that feeds our ML training and inference systems. You will design high-throughput, distributed data systems on top of columnar and lakehouse formats, so that GPUs and models are never starved for data. This is a hands-on role for someone who enjoys performance engineering and wants to work close to the boundary between data systems and machine learning.
Duties and Responsibilities:
- Design, build, and operate large-scale distributed data systems that serve ML training and inference workloads in production.
- Strong Python Engineer
- in Rust (or C++/Go) with Python bindings for ML practitioners to build high-performance data loading, storage, and retrieval layers (Systems Performance, which is critical work)
- Evaluate and adopt columnar and lakehouse formats (Parquet, Iceberg, Delta, Lance), and make the trade-offs explicit for schema evolution, random access, scan performance, and versioning.
- Optimize I/O-bound pipelines using Arrow, zero-copy techniques, memory mapping, async I/O, and efficient object storage access patterns (request coalescing, prefetching, caching, parallel range reads).
- Profile and remove bottlenecks across the data path, from object storage to host memory to accelerator, so training and inference stay compute-bound rather than I/O-bound.
- Partner with ML researchers and engineers to understand how training loops, evaluation, and inference services consume data, and turn that into system requirements.
- Define reliability, observability, and cost standards for data infrastructure, including SLOs, monitoring, capacity planning, and incident response.
- Write design documents, review code, and mentor engineers on performance engineering and distributed systems practices.
Requirements and Qualifications:
- Extensive experience building and running large-scale distributed data or ML infrastructure in production.
- Strong programming skills in Python, plus a systems language for performance-critical work (Rust strongly preferred; C++ or Go acceptable).
- Deep familiarity with columnar and lakehouse formats (Parquet, Iceberg, Delta, or Lance) and the trade-offs between them.
- Hands-on performance engineering for I/O-bound workloads: Arrow, zero-copy, memory mapping, async I/O, and high-throughput object storage access patterns.
- Working knowledge of the end-to-end ML workflow and how training and inference workloads consume data, enough to design data systems that serve them well.
Preferred Qualifications:
- Production Rust experience, including async runtimes (Tokio), FFI, and Python bindings (PyO3 or similar).
- Contributions to open-source data or ML infrastructure projects such as Arrow, Lance, Iceberg, Ray, or PyTorch data loading.
- Experience with distributed computing and orchestration frameworks (Spark, Ray, Dask, Kubernetes).
- Familiarity with multimodal or large-scale unstructured data (images, video, audio, embeddings) and vector or random-access storage.
- Experience with GPU-aware data pipelines, including pinned memory, GPUDirect Storage, and overlapping I/O with compute.
- Track record of cost optimization for cloud object storage and egress at petabyte scale.
- Experience leading technical design across teams and mentoring other engineers.
- BS/MS/PhD in Computer Science or a related field, or equivalent practical experience.
Bayside Solutions, Inc. is not able to sponsor any candidates at this time. Additionally, candidates for this position must qualify as a W2 candidate.
Bayside Solutions, Inc. may collect your personal information during the position application process. Please reference Bayside Solutions, Inc.'s CCPA Privacy Policy at