We are seeking a Staff MLOps Engineer to join an innovative Series B robotics and automotive technology startup in Woburn, MA. This is a full-time opportunity for a senior machine learning infrastructure engineer who wants to own the production lifecycle of computer vision models powering real-world automotive service applications. You'll work across Google Cloud Platform, Docker, Python, PyTorch/TensorFlow, model serving, CI/CD, infrastructure as code, and scalable inference systems.
This is a true ownership role where you'll take an existing machine learning product from MVP to production scale. You'll own the architecture behind how computer vision models are deployed, monitored, evaluated, improved, and served to customers. The technology is already being used in real automotive service environments, so your decisions will have a direct impact on product performance, customer experience, and the company's ability to scale. You'll also serve as a senior cloud architecture voice alongside the broader engineering team, making this a great opportunity for someone who wants significant technical ownership without stepping away from hands-on engineering.
Required Skills & Experience
8+ years of professional engineering experience
Several years of experience owning machine learning systems in production
Proven experience taking computer vision or ML pipelines from prototype through production scale
Deep experience with segmentation and classification models, preferably in multi-stage pipelines
Strong cloud infrastructure experience, preferably with Google Cloud Platform
Experience with Vertex AI, Cloud Run, GKE, Cloud Functions, Cloud SQL, Pub/Sub, and Docker
Strong Python development skills
Production experience with PyTorch or TensorFlow
Hands-on experience with model serving and optimization
Experience with technologies such as Triton, TorchServe, ONNX, TensorRT, quantization, or similar platforms
Experience owning production deployments, incident response, rollback procedures, and on-call support
Experience building data collection, labeling, and dataset versioning workflows
Experience developing reproducible model evaluation and regression testing systems
Experience with Terraform or similar infrastructure-as-code tools
Experience with CI/CD automation such as GitHub Actions
Strong understanding of the trade-offs between model accuracy, latency, infrastructure cost, and reliability
Excellent debugging, problem-solving, and communication skills
Desired Skills & Experience
Experience with on-device or edge inference
Experience with Core ML, TensorFlow Lite, or ExecuTorch
Experience building active learning or human-in-the-loop labeling systems
Computer vision experience with smaller, long-tail, or industrial inspection datasets
Experience integrating cameras, sensors, or other hardware with ML systems
Robotics, IoT, or edge computing experience
Experience with ROS or similar robotics platforms
Experience integrating ML capabilities into mobile applications
Automotive service, dealership, DMS, or automotive technology experience
Experience working in an early-stage or high-growth startup
Open-source contributions
Experience collaborating closely with hardware and field operations teams
What You Will Be Doing
Own the architecture and operation of the company's multi-stage computer vision inference pipeline
Re-architect the current MVP infrastructure into a scalable production serving platform
Design containerized inference infrastructure with GPU acceleration, queueing, batching, and autoscaling where appropriate
Own model deployment, versioning, staged rollouts, canary testing, shadow evaluation, and rollback
Build the data and model improvement lifecycle from field data collection through labeling, training, evaluation, and deployment
Establish dataset versioning and reproducible evaluation processes
Monitor model performance and identify drift or degradation in production
Analyze segmented model performance and investigate real-world failures
Define and monitor commercial ML metrics including false positives, false negatives, technician overrides, latency, and inference cost
Improve model performance through architecture selection, augmentation, hard-example mining, quantization, and distillation
Evaluate cloud versus on-device inference and help determine the right architecture
Build MLOps foundations including experiment tracking, reproducible training, model CI/CD, and infrastructure as code
Partner with senior software engineers to establish cloud architecture and Google Cloud Platform best practices
Work with hardware and field teams to improve image capture quality, including lighting, focus, and probe positioning
Own production support and incident response for the ML platform
Participate in on-call responsibilities for inference availability
Identify technical risks and communicate architectural trade-offs to engineering and company leadership
Tech Breakdown
25% MLOps / Model Deployment & Serving
20% Google Cloud Platform / Cloud Infrastructure
20% Computer Vision / ML Engineering
15% Data Pipelines / Evaluation / Model Improvement
10% DevOps / Infrastructure as Code / CI/CD
10% Architecture / Technical Leadership
Daily Responsibilities
45% Hands-On Engineering
20% ML Infrastructure & Production Operations
15% Model Performance & Evaluation
10% Architecture & Technical Leadership
10% Cross-Functional Collaboration
The Offer
Competitive salary
Comprehensive benefits package
Opportunity to own the ML platform for a product already being used by paying customers
Significant technical ownership over the company's production architecture
Opportunity to work with real-world proprietary computer vision data
Hands-on exposure to robotics, automotive technology, computer vision, and edge computing
Work alongside experienced robotics and software engineering professionals
Collaborative, low-ego, high-intensity startup environment
Prime Woburn, MA location with on-site parking
Opportunity to have a direct impact on the company's ability to scale
You will receive the following benefits:
Medical Insurance
Dental Benefits
Vision Benefits
Paid Time Off (PTO)
401(k)
Comprehensive Benefits Package
Professional Development Opportunities
Collaborative Startup Environment
On-Site Parking
Applicants must be currently authorized to work in the US on a full-time basis now and in the future.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 10105282
- Position Id: 891725
- Posted 13 hours ago