Sr. Software Engineer, ML Platform

Hybrid in San Francisco, CA, US • Posted 1 day ago • Updated 1 day ago
Full Time
No Travel Required
Hybrid
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🧠 Analyzing your skills...

Job Details

Skills

  • Python
  • SQL
  • Spark
  • PySpark
  • DataBricks
  • MLflow
  • AirFLow
  • AWS
  • SnowFlake
  • Kafka
  • Kinesis

Summary

Senior Software Engineer, ML Platform
Job Type: Full-time
Location: San Francisco, CA
Work Model: Hybrid
Experience: 5+ years


About the Position

We’re looking for a software engineer to join the Infrastructure team and lead the evolution of our ML Platform. This role is critical to building reliable, scalable, and developer-friendly systems for model experimentation, training, evaluation, inference, and retraining that power underwriting and other ML-driven products for small businesses.

As a Software Engineer, you’ll design, build, and maintain the core abstractions and platforms that let data scientists ship high-quality models to production—safely and quickly. You’ll partner closely with Data Science and Platform Engineering, own the ML platform end-to-end, and develop batch and real-time underwriting infrastructure.

 

What You''ll Do

  • Turn notebooks into software. Decompose data scientist training/inference notebooks into reusable, tested components (libraries, pipelines, templates) with clear interfaces and documentation.
  • Create developer-friendly ML abstractions. Build SDKs, CLIs, and templates that make it simple to define features, train/evaluate models, and deploy to batch or real-time targets with minimal boilerplate.
  • Build our real-time ML inference platform. Stand up and scale low-latency model serving.
  • Expand batch ML inference. Improve scheduling, parallelism, cost controls, observability, and failure/rollback for large-scale batch scoring and post-processing.
  • Own and expand the feature store. Design offline/online feature definitions, high read/write throughput, and consistent offline/online semantics.
  • Platform reliability and observability. Instrument training/inference for latency, throughput, accuracy, drift, data quality, and cost; build alerting and dashboards; drive incident response and postmortems.
  • Underwriting infrastructure partnership. Support production batch and real-time underwriting systems in collaboration with Data Science; collaborate on model interfaces, SLAs, safety checks, and product integrations.

What We Are Looking For

  • 5+ years of software engineering experience, including experience on ML platform/MLOps systems (training, deployment, and/or feature pipelines).
  • Strong Python; solid software design and testing fundamentals.
  • Proficiency with SQL.
  • Hands-on Spark/PySpark experience.
  • Knowledge of ML fundamentals—probability & statistics, supervised vs. unsupervised learning, bias/variance & regularization, feature engineering, model evaluation metrics, validation strategies, and production concerns like drift, stability, and monitoring.
  • Expertise with modern data/ML stacks—AWS, Databricks (workflows, lakehouse, MLflow/registry, Model Serving), and Airflow (or equivalent orchestration).
  • Experience building real-time systems (service design, caching, rate limiting, backpressure) and batch pipelines at scale.
  • Practical knowledge of feature-store concepts (offline/online stores, backfills, point-in-time correctness), model registries, experiment tracking, and evaluation frameworks.
  • Strong problem-solving skills and a proactive attitude toward ownership and platform health.
  • Excellent communication and collaboration skills, especially in cross-functional settings.

Bonus Points

  • Databricks experience (MLflow, Model Serving).
  • Experience with feature stores (e.g., Tecton, Feast) and streaming (Kafka/Kinesis).
  • Experience with fintech, risk, or underwriting systems; familiarity with model safety checks, rejection/override flows, and auditability.
  • Background with A/B testing platforms, shadow/canary deployments, and automated rollback.
  • Experience with low-latency inference systems.

Tech Stack

Python, SQL, Spark, PySpark, Databricks, MLflow, Airflow, AWS, Snowflake, Kafka, Kinesis

 

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91173258
  • Position Id: 9054236
  • Posted 1 day ago

Company Info

About THE TILTED CIRCLE LLC

Since 2017, we have provided international talent services to global conglomerates across multiple geographies. Our success is built on long-standing customer relationships and an elite clientele.

Contact the job poster
AP

Anand Pandey

Recruiter @ THE TILTED CIRCLE LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs