Senior Platform & Infrastructure Engineer MLBots

Hybrid in Rancho Park, CA, US • Posted 1 hour ago • Updated 1 hour ago
Contract Corp To Corp
Contract W2
6 Months
No Travel Required
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

🎯 Assessing qualifications...

Job Details

Skills

  • Amazon Web Services
  • Artificial Intelligence
  • Banking
  • Client/server
  • Cloud Computing
  • Computer Networking
  • Continuous Delivery
  • Continuous Integration
  • Dashboard
  • Data Collection
  • Energy
  • Evaluation
  • FOCUS
  • Financial Services
  • GPU
  • Good Clinical Practice
  • Google Cloud Platform
  • HPC
  • Health Care
  • Insurance
  • Kubernetes
  • LOS
  • Machine Learning (ML)
  • Machine Learning Operations (ML Ops)
  • Management
  • Mentorship
  • Microservices
  • Optimization
  • Orchestration
  • Pharmaceutics
  • Public Sector
  • Python
  • ROOT
  • Recruiting
  • Retail
  • Scheduling
  • Software Engineering
  • Technical Direction
  • Telecommunications
  • Training
  • Unreal Engine
  • Workflow

Summary

Senior Platform & Infrastructure Engineer MLBots
Hybrid Los Angeles, CA

As a Senior Platform & Infrastructure Engineer on the MLBots team, you will design, build, and operate the core infrastructure and ML platforms behind Riot's Game Understanding Agents. Your focus will be on the compute and orchestration platforms that power large-scale distributed training of Game Understanding Agents (e.g. via RL, IL, and other techniques), simulation environments, and policy evaluation, as well as the CI/CD, infrastructure-as-code, observability, and developer tooling that keep these systems production-grade. You will close critical infrastructure gaps across the team's stack, driving improvements to standards, automation, and operational maturity. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team.

We encourage you to apply even if you do not believe you meet every single qualification.

Responsibilities

  • Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation.
    Design infrastructure for running game simulation environments at scale, enabling parallel rollouts, data collection, training, and evaluation.
    Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments.
    Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity.
    Build observability, monitoring, alerting, health indicators, and SLO-aligned dashboards for infrastructure and ML workloads.
    Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows.
    Support MLOps workflows including automated training pipelines, model artifact management, experiment tracking, and reproducible ML lifecycle operations.
    Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles.

Qualifications

  • Bachelor s degree in Computer Science or a related field, or equivalent practical experience.
    3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles.
    Experience operating distributed systems in production and keeping them healthy under real load.
    Strong experience with Kubernetes, AWS or Google Cloud Platform, infrastructure-as-code, CI/CD, deployment automation, and production tooling.
    Experience with GPU compute infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running training workloads.
    Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals.
    Familiarity with MLOps workflows such as model versioning, pipeline orchestration, experiment tracking, artifact management, and reproducible ML workflows.
    Bonus: experience with distributed training or HPC frameworks, inference serving, systems languages, high-performance networking, game AI or simulation, Unreal/client-server architecture, AI-assisted development tools, or a passion for games and player experience.


About AgreeYa:
AgreeYa is a global systems integrator delivering a competitive advantage for its customers through software, solutions, and services. Established in 1999, AgreeYa is headquartered in Folsom, California, with a global footprint and a team of more than 1,800+ professionals across offices. AgreeYa works with 550+ organizations ranging from Fortune 100 firms to small and large businesses across industries such as Telecom, Banking, Financial Services & Insurance, Healthcare, Utility & Energy, Technology, Public Sector, Pharma & Biotech, Retail, Client, and others. Please visit us at for more information.
Equal Opportunity:
AgreeYa is an equal opportunity employer. We evaluate qualified applicants without regard to race, color, religion, gender identity, sexual orientation, national origin, disability, veteran status or other protected characteristics. Visit our website at to learn about our Career & Cultures.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: swapps
  • Position Id: 9056763
  • Posted 1 hour ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

West Hollywood, California

Today

Full-time

USD 190,000.00 - 246,000.00 per year

Los Angeles, California

Today

Full-time

USD 229,200.00 - 319,500.00 per year

Hawthorne, California

Today

Full-time

Burbank, California

Today

Full-time

USD 157,000.00 - 235,000.00 per year

Search all similar jobs