Senior MLOps Engineer
Location: 100% Remote
Duration: 5 Months Contract
Employment Type: W2 Only
Rate: Up to $90/hr (flexible for the right candidate)
Overview
We are seeking a Senior MLOps Engineer to help scale and operationalize machine learning platforms supporting predictive analytics and AI-driven applications. This role sits at the intersection of Data Science, DevOps, Cloud Infrastructure, and Platform Engineering, focused on taking models from experimentation through fully automated, monitored, and production-ready deployments.
The ideal candidate will have deep expertise in AWS SageMaker, Terraform, CI/CD automation, and ML platform operations, along with a strong understanding of machine learning concepts and model lifecycle management.
Required Qualifications
- 5+ years of experience in MLOps, ML Platform Engineering, or ML Infrastructure Engineering with ownership of production ML systems.
- Deep AWS expertise, including:
- SageMaker (training jobs, pipelines, model registry, endpoints)
- S3
- IAM
- KMS
- CloudWatch
- Lambda
- Step Functions
- Multi-account AWS environments
- Strong DevOps and Infrastructure-as-Code experience using:
- Terraform
- GitLab CI/CD, GitHub Actions, or similar tools
- Docker
- Git workflows
- Strong understanding of machine learning fundamentals, including:
- Model training and evaluation
- Feature engineering
- AUC, calibration metrics, C-index, and performance monitoring
- Advanced Python development experience with production-quality, well-tested code.
- Experience with model monitoring, drift detection, and model lifecycle management.
- Strong security knowledge, including least-privilege access controls, encryption, and secrets management.
Key Responsibilities
ML Platform & Lifecycle Management
- Design, build, and maintain end-to-end ML training and inference pipelines.
- Manage model registry, versioning, validation, promotion, and production deployments.
- Implement blue/green deployment strategies and automated rollback mechanisms.
Infrastructure as Code
- Develop and maintain Terraform modules supporting:
- SageMaker
- S3
- KMS
- IAM
- CloudWatch
- Cross-account AWS infrastructure
- Support development, staging, and production environments.
CI/CD Automation
- Design and maintain CI/CD pipelines covering:
- Automated testing
- Infrastructure security scanning
- Static code analysis
- Dependency validation
- Terraform plan/apply workflows
- Automated model promotion
Production Operations
- Build observability solutions using CloudWatch dashboards, metrics, logging, and alerting.
- Implement model and data drift detection strategies.
- Improve reliability, performance, and operational support for ML services.
Security & Governance
- Implement secure cross-account access patterns.
- Manage encryption, secrets, model artifact integrity, and compliance controls.
- Monitor infrastructure and AI workload costs.
AI & LLM Operations
- Support deployment and management of LLM-powered workloads.
- Optimize throughput, cost efficiency, monitoring, and operational guardrails for generative AI services.
Data Science Partnership
- Collaborate closely with Data Scientists to productionize experimental models.
- Establish reusable MLOps standards, best practices, and deployment patterns.
AIOps & Automation
- Implement anomaly detection across model, infrastructure, and cost signals.
- Design automated remediation workflows, including scaling and rollback mechanisms.
- Leverage AI-assisted observability and incident analysis to reduce operational overhead and improve MTTR.
Preferred Experience
- AWS Bedrock
- XGBoost and predictive modeling workloads
- Model governance and ML platform standardization
- AIOps and intelligent monitoring solutions
- Enterprise-scale, multi-account AWS environments
- Healthcare, clinical, or regulated industry experience
Top Skills
- AWS SageMaker
- Terraform
- Python
- GitLab CI/CD
- MLOps
- Model Monitoring & Drift Detection
- Docker
- IAM/KMS Security
- CloudWatch
- AWS Bedrock
Keywords: MLOps Engineer, ML Platform Engineer, SageMaker Engineer, AWS AI Engineer, Machine Learning Infrastructure Engineer, Terraform, CI/CD, Python, Bedrock, Model Deployment, Data Science Platform, ML Operations.