Title- Senior ML Engineer Platform & Integrations
Location- Los Angeles ,CA Hybrid (3 Days Onsite/Week)
Duration 12 Month (Contract to Start)
Job Description-
As a Senior ML Engineer on the MLBots team, you will design, build, and operate the ML systems behind Game Understanding Agents. These agents learn to play and understand our games through reinforcement learning, imitation learning, and other techniques. You will build distributed training systems, simulation integration, evaluation pipelines, serving infrastructure, and developer tooling that enable the team to train and deploy agents at scale. You are a generalist, comfortable moving between training code, evaluation infrastructure, and production systems depending on where the biggest gaps are. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team.
We encourage you to apply even if you do not believe you meet every single qualification.
Responsibilities
Build distributed training systems for Game Understanding Agents across multi-node GPU clusters, integrating with RL/IL frameworks and optimizing for training throughput and convergence.
Design and operate game simulation environments at scale, enabling parallel rollouts, experience collection, and policy evaluation.
Build automated, reproducible ML pipelines that version data and model artifacts, validate agents against quality thresholds, and track experiment lineage.
Build serving systems for agent inference, including model version routing, fallback behavior, and optimization across latency, cost, and accuracy.
Build developer-facing tools, SDKs, and templates that simplify the agent development loop, from local prototyping to large-scale distributed training and evaluation. Validate adoption through developer feedback.
Define health indicators, drift detection metrics, training diagnostics, and alerting aligned with SLOs. Create comprehensive model and system documentation.
Identify, manage, and resolve production incidents including non-deterministic failure modes such as training instability, reward signal degradation, and policy regression.
Build governance and security controls in partnership with compliance, security, and privacy teams. Ensure model lifecycle compliance and audit trails.
Mentor less experienced ML engineers in craft skills and technical best practices.
Participate in recruiting for ML engineering roles. Contribute to interview loops and candidate evaluation.
Qualifications
Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
3+ years of experience in software engineering, with substantial time in ML engineering, platform, or systems roles.
Track record operating distributed systems in production, not just deploying them, but keeping them healthy under real load.
Experience with distributed training of ML agents, including familiarity with RL, IL, or other agent training paradigms at scale.
Proficiency in Python, working knowledge of C++, and willingness to engage with game engine technologies (e.g., Unreal Engine, Blueprints, Lua). Hands-on Experience with ML frameworks (e.g., PyTorch) and familiarity with distributed RL/IL training frameworks (e.g., Ray/RLlib).
Experience with GPU-accelerated workloads, including multi-node orchestration and resource optimization for long-running training.
Comfort working across the stack, able to move between training algorithms, evaluation systems, and production infrastructure depending on where the team needs you.
Strong written communication. Able to document systems clearly and contribute to team-wide standards.
Experience with game simulation, game AI, or real-time interactive environments is a plus.
Background in inference serving, model evaluation frameworks, or ML governance in production systems is a plus.
Prior success standing up greenfield platform work or building toward self-service ML capabilities is a plus.
Passion for player experience, games, or creative technology.