TRACE brings FP4 rollouts to MoE reinforcement learning, up to 5.4x faster

Reinforcement learning is a standard way to post-train large language models, but it is costly. The authors point to the rollout stage, where the model generates its own samples, as a source of substantial computation and memory overhead. That is why low-precision rollout is attractive for efficient RL training.
The paper's complaint about existing FP4 RL methods is specific. They mostly optimize quantization accuracy on the training path and the rollout path independently. They do not directly reduce the discrepancy between the two quantized execution paths, and the authors call this a key limitation.
Their answer is TRACE, short for Train-Rollout Quantization Alignment via Compact GuidancE. It is an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models. It has two parts. The first is rollout-guided quantization-aware training: quantization outcomes seen on the rollout side are used to guide FP4 rounding decisions on the training side, which directly narrows the train-rollout discrepancy. The second is a quantization-information caching scheme. Passing rollout guidance to the trainer adds storage and communication overhead, so the scheme selectively retains only mantissa and scale information from deeper layers to keep that overhead down.
The authors evaluated TRACE on four large-scale MoE language models, across reasoning, coding and long-horizon RL tasks. They report that it enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout. The same results show up to 5.4x rollout speedup, and strong final FP4 performance compared with post-hoc FP4 quantization of policies that were trained in BF16.
Key facts
- TRACE (Train-Rollout Quantization Alignment via Compact GuidancE) is an FP4 quantization framework for RL training of Mixture-of-Experts language models.
- It uses rollout-side quantization outcomes to guide training-side FP4 rounding, directly reducing the discrepancy between the training and rollout paths.
- A caching scheme keeps only mantissa and scale information from deeper layers, cutting the storage and communication cost of that guidance.
- Evaluated on four large-scale MoE models across reasoning, coding and long-horizon RL tasks, with up to 5.4x rollout speedup.
- The authors report RL performance comparable to BF16 rollout, with joint FP4 weight/activation and FP4 KV-cache rollout.
Why it matters
Rollout generation is a heavy part of RL post-training for LLMs, in both computation and memory. Running it in FP4 is an obvious way to save on both, but the authors say existing FP4 RL methods tune the training and rollout paths separately and leave the gap between them in place. TRACE targets that gap directly, and the reported result is FP4 rollout that keeps RL performance comparable to BF16 rollout while running up to 5.4x faster.
Who it affects
The work is aimed at teams doing RL post-training of Mixture-of-Experts language models, where rollout cost is the problem being addressed. It is a research result, so the nearest audience is researchers and engineers working on low-precision training and RL infrastructure.
How to use it
The abstract describes a method and does not say how to apply it. The source does not mention a code or model release. Practitioners can take away the design idea: let rollout-side quantization outcomes steer training-side FP4 rounding, and cache only mantissa and scale information from deeper layers to limit overhead.
How solid is it
The claims come from the paper's abstract, so these are the authors' own reported results. They cover four large-scale MoE models and three kinds of task: reasoning, coding and long-horizon RL. The 5.4x figure is a maximum ("up to"), and the baseline for it is not stated. "Comparable to BF16 rollout" is not quantified. The four models are not named, and no benchmark scores are given.
Risks and caveats
The headline speedup is a best case, not a typical value, and its baseline is not stated. The abstract names no benchmarks, accuracy numbers, hardware or model sizes, so the size of any quality gap against BF16 cannot be judged from it. The abstract also names no authors or institutions. Independent replication would be needed before treating the results as general.