MiMo-V2.6 report details scaling RL compute along three dimensions

MiMo-V2.6 report details scaling RL compute along three dimensions

A technical report introduces the MiMo-V2.6 series, an omni-modal family of models that, in the authors' words, pushes the frontier of model intelligence by scaling reinforcement learning compute. The report starts from the premise that RL is the central training paradigm for advancing large foundation models towards self-improvement.

Before RL begins, the team runs mid-training on a broad multimodal corpus, which they say provides ample exploration space. They also build infrastructure on the pretrained hybrid-SWA architecture to support the later scale-up.

RL compute is then scaled along three dimensions. First, larger batches and higher throughput: the asynchronous training consumes 1,568 samples and 2.7-3.7B tokens per step, at context lengths of up to 1M. Second, more diverse and complex environments, spanning code, general, visual and cyber domains under a mixture of agent harnesses. Third, more grader compute, through groupwise agentic grading, which the authors say yields more accurate reward signals for long-horizon tasks and steers the model towards shorter, more token-efficient solutions.

Two further pieces address stability and plumbing. To keep training stable at scale, the team freezes the MoE router and establishes a multi-layer defense against reward hacking. For mixed-task agentic RL they build infrastructure that includes a unified trajectory representation, high-concurrency multi-framework rollout, decoupled control and data planes, and training-inference consistency.

Finally, the authors say they open-source the training dynamics, RL environments and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.

Key facts

  • MiMo-V2.6 is an omni-modal model family whose gains come from scaling RL compute, after mid-training on a broad multimodal corpus.
  • RL compute is scaled along three dimensions: batch size and throughput, environment diversity, and grader compute.
  • The asynchronous RL training consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M.
  • Stability measures include freezing the MoE router and a multi-layer defense against reward hacking.
  • The authors say the training dynamics, RL environments and RL framework are open-sourced.

Why it matters

The report treats reinforcement learning as the central paradigm for pushing large foundation models towards self-improvement, and it lays out concrete numbers for what scaling RL looks like in practice: 1,568 samples and 2.7-3.7B tokens per step, with context up to 1M. It also names the practical problems that come with that scale, namely router instability and reward hacking, and says how the team handled them. Publishing the training dynamics, environments and framework is the part that goes beyond a model announcement.

Who it affects

Mainly researchers and engineering teams working on RL for large models, especially agentic and multimodal setups. The environments span code, general, visual and cyber domains, so groups building agent training pipelines in those areas are the closest audience. Readers who only want a model to use will find little here beyond the description of how it was trained.

How to use it

The authors say the training dynamics, RL environments and RL framework are open-sourced to support reproduction and further research on scaled RL and model self-improvement. The practical takeaways are the design choices they describe: groupwise agentic grading for long-horizon tasks, freezing the MoE router for stability, a multi-layer defense against reward hacking, and a unified trajectory representation with decoupled control and data planes. A release date, license or repository link is not stated.

How solid is it

This is a technical report, and the account here rests on its abstract, where the claims are the authors' own. The abstract names no authors or institutions. No benchmark results, scores or comparisons with other models are given, so the claim of pushing the frontier of model intelligence is not backed by figures in the text available. No parameter counts, model sizes or MoE configuration details are given either. The abstract does not state the total RL compute or training duration.

Risks and caveats

The abstract does not say that self-improvement was achieved, only that RL is the paradigm toward it. The reported throughput figures describe training consumption, not model quality. Reward hacking is named as a real hazard at this scale, and the multi-layer defense is described only in a single sentence. Reproduction claims depend on the open-sourced material actually being available and complete, which the abstract does not let a reader confirm.

“We open-source the training dynamics, RL environments, and RL framework to facilitate reproduction and further research on scaled RL and model self-improvement.”

— MiMo-V2.6 technical report abstract