Senior MLOps Engineer (AWS SageMaker, Terraform, Python)

Remote • Posted 1 hour ago • Updated 1 hour ago
Contract W2
6 Months
No Travel Required
Remote
$80 - $90/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Amazon Web Services
  • Amazon SageMaker
  • DevOps
  • Machine Learning Operations (ML Ops)
  • Python
  • Data Science

Summary

Senior MLOps Engineer

Location: 100% Remote
Duration: 5 Months Contract
Employment Type: W2 Only
Rate: Up to $90/hr (flexible for the right candidate)

Overview

We are seeking a Senior MLOps Engineer to help scale and operationalize machine learning platforms supporting predictive analytics and AI-driven applications. This role sits at the intersection of Data Science, DevOps, Cloud Infrastructure, and Platform Engineering, focused on taking models from experimentation through fully automated, monitored, and production-ready deployments.

The ideal candidate will have deep expertise in AWS SageMaker, Terraform, CI/CD automation, and ML platform operations, along with a strong understanding of machine learning concepts and model lifecycle management.


Required Qualifications

  • 5+ years of experience in MLOps, ML Platform Engineering, or ML Infrastructure Engineering with ownership of production ML systems.
  • Deep AWS expertise, including:
    • SageMaker (training jobs, pipelines, model registry, endpoints)
    • S3
    • IAM
    • KMS
    • CloudWatch
    • Lambda
    • Step Functions
    • Multi-account AWS environments
  • Strong DevOps and Infrastructure-as-Code experience using:
    • Terraform
    • GitLab CI/CD, GitHub Actions, or similar tools
    • Docker
    • Git workflows
  • Strong understanding of machine learning fundamentals, including:
    • Model training and evaluation
    • Feature engineering
    • AUC, calibration metrics, C-index, and performance monitoring
  • Advanced Python development experience with production-quality, well-tested code.
  • Experience with model monitoring, drift detection, and model lifecycle management.
  • Strong security knowledge, including least-privilege access controls, encryption, and secrets management.

Key Responsibilities

ML Platform & Lifecycle Management

  • Design, build, and maintain end-to-end ML training and inference pipelines.
  • Manage model registry, versioning, validation, promotion, and production deployments.
  • Implement blue/green deployment strategies and automated rollback mechanisms.

Infrastructure as Code

  • Develop and maintain Terraform modules supporting:
    • SageMaker
    • S3
    • KMS
    • IAM
    • CloudWatch
    • Cross-account AWS infrastructure
  • Support development, staging, and production environments.

CI/CD Automation

  • Design and maintain CI/CD pipelines covering:
    • Automated testing
    • Infrastructure security scanning
    • Static code analysis
    • Dependency validation
    • Terraform plan/apply workflows
    • Automated model promotion

Production Operations

  • Build observability solutions using CloudWatch dashboards, metrics, logging, and alerting.
  • Implement model and data drift detection strategies.
  • Improve reliability, performance, and operational support for ML services.

Security & Governance

  • Implement secure cross-account access patterns.
  • Manage encryption, secrets, model artifact integrity, and compliance controls.
  • Monitor infrastructure and AI workload costs.

AI & LLM Operations

  • Support deployment and management of LLM-powered workloads.
  • Optimize throughput, cost efficiency, monitoring, and operational guardrails for generative AI services.

Data Science Partnership

  • Collaborate closely with Data Scientists to productionize experimental models.
  • Establish reusable MLOps standards, best practices, and deployment patterns.

AIOps & Automation

  • Implement anomaly detection across model, infrastructure, and cost signals.
  • Design automated remediation workflows, including scaling and rollback mechanisms.
  • Leverage AI-assisted observability and incident analysis to reduce operational overhead and improve MTTR.

Preferred Experience

  • AWS Bedrock
  • XGBoost and predictive modeling workloads
  • Model governance and ML platform standardization
  • AIOps and intelligent monitoring solutions
  • Enterprise-scale, multi-account AWS environments
  • Healthcare, clinical, or regulated industry experience

Top Skills

  • AWS SageMaker
  • Terraform
  • Python
  • GitLab CI/CD
  • MLOps
  • Model Monitoring & Drift Detection
  • Docker
  • IAM/KMS Security
  • CloudWatch
  • AWS Bedrock

Keywords: MLOps Engineer, ML Platform Engineer, SageMaker Engineer, AWS AI Engineer, Machine Learning Infrastructure Engineer, Terraform, CI/CD, Python, Bedrock, Model Deployment, Data Science Platform, ML Operations.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91172532
  • Position Id: 9085506
  • Posted 1 hour ago
Contact the job poster
RS

Ritesh Sharma

Recruiter @ SCMInnovators LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote or Chicago, Illinois

Today

Easy Apply

Contract

$50 - $80

Remote or Atlanta, Georgia

Today

Contract

$40 - $50 hourly

Remote

Today

Easy Apply

Contract

$60 - $70

Remote or Chicago, Illinois

Today

Easy Apply

Full-time

$100000 - $120000

Search all similar jobs