MLOPS Platform Engineer

Remote • Posted 9 hours ago • Updated 9 hours ago
Contract Independent
Contract W2
Contract Corp To Corp
12 Months
No Travel Required
Able to Sponsor
Remote
$70 - $75/hr
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • MLOPS
  • Machine Learning Operations
  • AWS
  • Devops Cloud Infrastructure
  • Cloudwatch
  • VPC
  • Data Platform
  • tensorflow
  • Pytorch

Summary

Role: ML Ops & Data Platform Engineer

Location: USA (remote is good but need to cover PST time)

Duration: Long term project

 

ML Ops & Data Platform Engineer — Contractor Overview We are seeking an experienced ML Ops & Data Platform Engineer to join a team for projects covering two areas of focus: 1. ML Operations — help design, build, and operate production-grade machine learning infrastructure on AWS for advanced cancer screening and precision oncology applications, and support the integration of AI/ML into existing and new workloads. 2. Data Platforming — contribute to the establishment and creation of a shared Data Platform for use across the broader Science Office: a governed, self-service data foundation (ingestion, lakehouse architecture, cataloging, quality, and access control) serving multiple science teams and use cases, of which ML workloads are one consumer among many. These two workstreams are complementary, not parallel silos — the data platform is the foundation the ML pipelines will increasingly consume from. This role is platform-first: as a senior hands-on contributor working alongside internal data science, engineering, and broader Science Office teams, this person will drive key pieces of the AWS architecture, automation, and operational reliability of both the ML pipelines and the underlying data platform, sharing accountability for these systems with the rest of the team. Deliverables include working infrastructure-as-code, CI/CD pipelines, data ingestion pipelines, catalog/governance set up, observability, and documentation/knowledge transfer to internal teams as engagements close.

 

 

Job Duties Includes, but is not limited to, the following:

ML Operations Design, build, and operate AWS-native ML infrastructure, as part of a team, supporting batch and real-time inference workloads, including: Amazon SageMaker (training jobs, endpoints, pipelines, model registry) ECS/EKS + Docker for containerized services

AWS Lambda and Step Functions for orchestration and event-driven workflows S3 as the primary data/artifact layer, with lifecycle and access policies CloudWatch (and related tooling) for logging, metrics, and alerting IAM & VPC design following least-privilege and network security best practices Infrastructure as Code via AWS CDK or CloudFormation Implement CI/CD pipelines for ML (data, model, and code), including automated testing, packaging, and promotion of models across dev/staging/production environments. Help establish model and data versioning, experiment tracking, and lineage for ML pipelines to support reproducibility and auditability (builds on the shared platform''s catalog and lineage rather than duplicating it). Build monitoring, logging, and alerting for model performance, model-input data drift, and system health; help define SLOs/SLAs for critical ML services and build the automation needed to meet them. Collaborate with data scientists and software/platform engineering teams to translate experimental workflows into production-grade services.

 

Data Platform Engineering Contribute to the architecture and hands-on build of an S3-based data lake / lakehouse for the Science Office, including zone design (raw/curated/consumption), open table formats (e.g., Apache Iceberg), partitioning, and lifecycle policies. Build batch and streaming ingestion/ETL/ELT pipelines using AWS Glue (jobs, crawlers) Implement data cataloging, governance, and access control, in collaboration with governance/security stakeholders, using the Glue Data Catalog and Lake Formation (finegrained, cross-team permissions), and optionally DataZone for data product publishing/discovery across science teams Enable self-service analytics and data access for scientists across the Science Office via Athena including query patterns, cost controls, and onboarding documentation. Help establish data quality, schema/metadata management, and lineage frameworks (e.g., Glue Data Quality, schema registries) for shared datasets — distinct from ML-side model-input drift monitoring, this workstream covers dataset-level quality for assets consumed across multiple teams.

 

Shared Duties Identify and implement cost optimizations across both ML and data platform workloads (right-sizing, spot/scheduling strategies, storage tiering, Athena/Redshift query cost controls) without compromising reliability. Collaborate with data scientists, analysts, and scientists across the broader Science Office to translate experimental workflows and analytical needs into production-grade services. Document architecture, runbooks, and operational procedures for both the ML infrastructure and the data platform, and provide knowledge transfer to internal teams as part of engagement close-out.

 

 Required Skills:

Required Experience 5+ years of hands-on experience in MLOps, data engineering, DevOps, or cloud infrastructure engineering roles, with a strong recent focus on AWS.

Demonstrated production experience with a substantial subset of ML infrastructure services: SageMaker, ECS/EKS, Lambda, Step Functions, S3, CloudWatch, IAM, VPC, and CDK or CloudFormation for infrastructure as code. Demonstrated production experience with a substantial subset of AWS data platform services: Glue, Lake Formation, Athena.

Experience designing and operating data lake/lakehouse architectures on S3, including ETL/ELT pipeline development and open table formats (e.g., Iceberg). Experience with data governance, cataloging, and multi-team access control (Lake Formation, Glue Data Catalog, or equivalent) for shared, regulated data assets.

Strong Python and SQL skills, with working familiarity with at least one ML framework (e.g., TensorFlow, PyTorch, scikit-learn) sufficient to integrate with and support data science workflows. Experience with containerization and orchestration (Docker, Kubernetes) in production settings. Experience building and operating CI/CD pipelines for ML and data systems (data, model, and code promotion).

Must be authorized to work in the country where work will be performed.

 

Nice to Have AWS certifications (Solutions Architect – Professional, Machine Learning – Specialty, Data Analytics/Data Engineer) or equivalent demonstrated expertise. Experience with multi-account AWS architectures and GPU workload management on AWS. Experience with MLflow, SageMaker Model Registry, or similar model governance tooling. Experience with DataZone or data-mesh/data-product patterns. Experience with dbt, Airflow/MWAA, or similar transformation/orchestration tooling.

Experience standing up self-service data platforms consumed by non-engineering scientific/analytical users. Terraform experience as an alternative/complement to CDK. Prior experience delivering as a contractor/consultant with clean documentation and handoff practices. Prior experience in healthcare, life sciences, or other regulated environments.

 

Success Criteria / Definition of Done Success is measured by this contractor''s contributions to the following team outcomes: ML Operations Contractor-delivered ML pipeline components running in production with automated CI/CD. Monitoring, logging, and alerting live and validated against SLOs defined with the team. Data Platforming Data platform foundation live on AWS: S3 lakehouse zones defined, Glue Data Catalog populated, and Lake Formation permissions implemented for at least the initial Science Office teams/use cases. Ingestion pipelines operational for the agreed initial data sources, with data quality checks in place. Self-service query access (Athena/Redshift) validated by at least one non-ML science team, with onboarding documentation. Shared All infrastructure delivered as part of this engagement codified (CDK/CloudFormation) and checked into version control. Documentation and knowledge transfer completed with the internal team, covering both workstreams



Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91131106
  • Position Id: 9063846
  • Posted 9 hours ago

Company Info

About Rivago infotech inc

Rivago Infotech Inc has been a leader in IT staffing and Software development for over 5 years and is one of the largest diversity and development firms in the industry. We are known for our high-touch, customer-eccentric approach, offering our clients unmatched quality, responsiveness and flexibility . We are appreciated by our clients for our streamlined execution, highly efficient service and exceptional talent management that go above and beyond traditional staffing services.

About_Company_OneAbout_Company_Two
Contact the job poster
RA

Rajat Arora

Recruiter @ Rivago infotech inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs