Hello ,
My name is Charles Powell and I am a staffing Specialist at Voto Consulting LLC. I am reaching out to you on an exciting job opportunity with one of our clients.
Job Title : Software Engineer - ML Infrastructure
Location : Onsite - San Francisco, CA (5 Days a week)
Duration : Full Time Position
Visa : Any Visa
The Opportunity
We are looking for an ML Infrastructure Engineer with 6+ years of experience to come in and establish best practices for our distributed training, reinforcement learning, and inference systems. You'll be a senior, autonomous engineer who can teach the team what world-class ML infrastructure looks like — someone who's operated at scale at a top-tier engineering organization and is ready to own the systems that power our foundation models. This is a ground-floor opportunity at a 5-person team backed by Bedrock Capital, Trust Ventures, and Jack Altman, building the first radiology-as-a-service platform powered by a multi-modal AI copilot to solve a crisis that leads to 200k deaths per year.
What You'll Be Doing
- Own and build the distributed training infrastructure for foundation models on large-scale medical imaging, including parallelism and checkpointing for volumetric data
- Design and operate the reinforcement learning training stack — high-throughput rollout generation, reward-model serving, and experience collection at scale
- Partner directly with researchers (ex-DeepMind, ex-Meta) to understand their workflows, prototype new ideas, and translate them into production-ready systems
- Build high-throughput data loading and preprocessing pipelines that keep GPUs saturated on large volumetric and multimodal datasets
- Contribute to production serving and deployment pipelines, including model rollout, canary deployments, and monitoring
- 6+ years of experience as an ML infrastructure engineer or distributed systems engineer
Application Questions :
- Are you able to work in San Francisco and come into the office 5 days per week?
- Walk us through your experience building distributed training infrastructure (e.g. FSDP, DeepSpeed, or Megatron-style parallelism) and the scale you've operated at.
- Describe your experience building ML infrastructure or platforms that research teams rely on for experimentation and training workflows.
Thanks and Regards
Charles Powell || Lead Technical Recruiter
Voto Consulting LLC
Direct: (551)–274-5551
1549 Finnegan Lane, 2nd Floor, North Brunswick, NJ 08902
Error! Hyperlink reference not valid.