Lead Platform Engineer

Hybrid in South San Francisco, CA, US • Posted 3 hours ago • Updated 3 hours ago
Full Time
No Travel Required
Hybrid
$205000 - $235000/yr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Staff Backend
  • Platform Engineer
  • Python
  • Distributed Systems
  • Backend Engineering
  • Terraform
  • Docker
  • Asynchronous Processing
  • Orchestration
  • RabbitMQ
  • ECS

Summary

Lead Platform Engineer

Location: South San Francisco / San Mateo, CA
Work Arrangement: Hybrid — 3 days onsite per week
Employment Type: Full-Time
Openings: 2

Position Overview

We are seeking two hands-on Lead Platform Engineers to help build and scale the backend and cloud infrastructure supporting a global DNA-sequencing operation.

This is a backend/platform engineering leadership role, not a traditional DevOps, IT infrastructure, data engineering, bioinformatics, or ML model-training position.

The ideal candidate is a Senior or Staff-level engineer with strong Python, AWS, distributed-systems, and cloud-orchestration experience, who has worked in an early-stage startup environment and personally built systems from the ground up.

The role will initially be highly hands-on, with the engineer learning the existing systems and taking ownership of major technical areas. Over time, the position will transition into approximately:

  • 70% hands-on architecture and software engineering
  • 30% technical leadership, mentoring, and team leadership

What You'll Build

The platform supports a globally distributed sequencing operation involving:

  • Laboratory robots and DNA sequencers
  • On-premises Linux infrastructure
  • AWS cloud environments
  • Terabytes of sequencing data movement
  • Hundreds of thousands to millions of bioinformatics jobs
  • Distributed queues and asynchronous workloads
  • Workflow orchestration and scheduling
  • Retry and failure-handling mechanisms
  • Distributed state management
  • Fault-tolerant production systems
  • Large-scale compute infrastructure
  • AI agents that interact with internal tools, APIs, operational data, and physical workflows

The infrastructure currently supports approximately 5,000 CPUs, 12.5 TB of RAM, and 100+ GPUs, with significant production AWS usage.


Key Responsibilities

  • Design, build, and operate high-scale distributed backend and platform systems.
  • Develop production backend services primarily using Python.
  • Build and maintain cloud-native services and infrastructure on AWS.
  • Design systems for asynchronous processing, distributed queues, workflow orchestration, scheduling, retries, state management, and fault tolerance.
  • Build reliable systems for moving large volumes of data between laboratory equipment, on-premises infrastructure, and AWS.
  • Develop services that orchestrate large numbers of compute and bioinformatics workloads.
  • Design and implement scalable APIs, microservices, workers, and event-driven services.
  • Work across application code, cloud infrastructure, Linux systems, data movement, and operational tooling.
  • Architect and implement production systems from concept through deployment and ongoing operation.
  • Troubleshoot complex production issues and improve system reliability, performance, and scalability.
  • Lead major technical projects from design through production.
  • Establish engineering patterns, standards, and best practices for a growing platform team.
  • Mentor and provide technical guidance to engineers while remaining deeply hands-on.
  • Collaborate with engineering, infrastructure, security, and scientific teams.
  • Explore and implement practical applications of AI agents and modern AI development tools within production systems.

Required Qualifications

  • 6+ years of professional experience in backend engineering, platform engineering, infrastructure engineering, distributed systems, or closely related software engineering roles.
  • Strong professional experience with Python backend development.
  • Strong hands-on experience with AWS cloud infrastructure and services.
  • Significant experience designing and building distributed production systems.
  • Hands-on experience with task queues, asynchronous processing, event-driven systems, workflow orchestration, scheduling, retries, state management, and fault tolerance.
  • Experience building systems from scratch or from early-stage prototypes through production.
  • Meaningful experience working in an early-stage startup or rapidly scaling engineering environment.
  • Demonstrated ownership of major technical systems or projects.
  • Experience scaling systems, infrastructure, workloads, or engineering platforms as a company grows.
  • Experience providing technical leadership and mentoring engineers.
  • Strong Linux and cloud-infrastructure fundamentals.
  • Ability and willingness to remain approximately 70% hands-on with architecture, coding, debugging, and production engineering.

Required Technical Skills

Candidates should have strong experience with the following core technologies and concepts:

Backend & Programming

  • Python
  • FastAPI, Django, or comparable Python backend frameworks
  • REST APIs
  • Microservices
  • Event-driven architecture
  • Asynchronous processing
  • Backend service development

AWS & Cloud

  • Amazon Web Services (AWS)
  • ECS
  • AWS Batch
  • AWS Step Functions
  • SQS
  • Lambda
  • Cloud-native architecture
  • Infrastructure as Code
  • Production cloud environments

Distributed Systems & Orchestration

  • Distributed systems
  • Distributed computing
  • Task queues
  • Job queues
  • Message queues
  • Workflow orchestration
  • Job orchestration
  • Scheduling
  • Retry mechanisms
  • State management
  • Fault tolerance
  • Failure recovery
  • High-volume workload processing
  • Event-driven services

Messaging & Distributed Processing

Experience with one or more of:

  • Celery
  • Amazon SQS
  • Kafka
  • RabbitMQ
  • gRPC
  • Similar distributed messaging or task-processing technologies

Infrastructure

  • Linux
  • Docker / containerized environments
  • Cloud infrastructure
  • Infrastructure automation
  • Production monitoring and troubleshooting
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10504943
  • Position Id: 4163-12901-1787671294
  • Posted 3 hours ago

Company Info

About DKMRBH Inc.

Do Know Me Right Before Hiring knows the value of transparency & constant communication between the client & the project resources to build a trustworthy foundation of our name. A commitment made is a commitment honored. Our clients experience the commitment during our intensive project delivery engagements. We believe in building a reputation that will serve as a positive reflection on our clients & candidates. We measure our success by the strength of the relationships we cultivate.



At DKMRBH Inc. we understand the importance of work to our clients & aspire to honor all of our commitments to them. Deep technical & functional expertise in every aspect of the Enterprise management: Our senior management team brings a wealth of experience in leading companies, delivering a consistent record of positive results. This stems directly from our devotion to sustain, mutually benefit & build long-term relationships with our clients,

Contact the job poster
BK

Bisma Kahn

Recruiter @ DKMRBH Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs