TRIAGE stabilizes native NVFP4 reinforcement learning for LLMs

Running reinforcement learning (RL) for large language models in low precision can speed it up a lot, but there is a catch. The sampler, which generates responses, and the learner, which updates the policy, can execute slightly differently, and those discrepancies can destabilize policy optimization. A paper on Hugging Face Papers takes this problem on for native NVFP4, a 4-bit format, and proposes a fix called TRIAGE.
The authors start by studying how mismatch interacts with the policy-gradient direction. Their point is that some update contributions are locally amplifying and others are contracting, and that mismatch magnitude alone cannot tell the two apart. In native NVFP4 runs they observe an early imbalance between the two amplifying regions, favoring negative-advantage, negative-gap updates. The tail tokens of those updates become concentrated in a small fraction of response segments, and only after that does the mismatch spread globally.
TRIAGE is built on those observations. It is a direction-aware stabilization method with two parts: segment-level diagnosis that selectively rebalances policy-gradient updates, and a bounded repair applied to residual severe mismatch. It changes the optimization objective, not the numerics. The forward execution stays in native NVFP4 weight-and-activation 4-bit (W4A4) on both the sampler and the learner.
The experiments use Qwen3-4B and Qwen3-30B-A3B. The authors report stable optimization throughout the evaluated training horizon and full-precision-level performance across five mathematical reasoning benchmarks. Native NVFP4 with TRIAGE gives up to 2.3x higher rollout throughput than BF16.
Key facts
- TRIAGE is a direction-aware stabilization method for reinforcement learning of LLMs run natively in NVFP4, with W4A4 forward execution kept on both sampler and learner.
- The authors find that in native NVFP4 runs there is an early imbalance favoring negative-advantage, negative-gap updates, whose tail tokens concentrate in a small fraction of response segments before mismatch spreads globally.
- TRIAGE uses segment-level diagnosis to selectively rebalance policy-gradient updates and applies bounded repair to residual severe mismatch, changing the optimization objective only.
- Experiments on Qwen3-4B and Qwen3-30B-A3B show stable optimization over the evaluated training horizon and full-precision-level results on five mathematical reasoning benchmarks.
- Native NVFP4 with TRIAGE provides up to 2.3x higher rollout throughput than BF16.
Why it matters
RL is a heavy step in training reasoning models, and rollouts (generating responses) are a large part of its cost. Low precision can speed this up, but the paper says mismatch between learner and sampler execution can destabilize policy optimization. TRIAGE targets that instability while keeping 4-bit forward execution on both sides, which is where the reported speedup of up to 2.3x over BF16 comes from. The diagnosis is also a contribution: the authors argue that the direction of an update matters, not just the size of the mismatch.
Who it affects
Mainly teams that run RL on LLMs and want to cut the cost of rollouts with NVFP4. The evidence is on Qwen3-4B and Qwen3-30B-A3B and on mathematical reasoning, so the clearest relevance is to reasoning-focused RL on models of that kind.
How to use it
TRIAGE is a change to the optimization objective, not to the forward pass: native NVFP4 W4A4 execution stays on both the sampler and the learner. In practice it adds segment-level diagnosis, selective rebalancing of policy-gradient updates, and bounded repair of residual severe mismatch. The source gives no code release, license or availability information.
How solid is it
The claims come from the paper's own abstract and are the authors' results. They report stable optimization on two Qwen3 models and full-precision-level performance on five mathematical reasoning benchmarks. No authors or institutions are named in the source, and no per-benchmark scores or numeric accuracy figures are given. The source does not say which five benchmarks were used. No comparison with other stabilization methods is mentioned.
Risks and caveats
The 2.3x figure is a maximum ('up to'); the source does not say on which model or setting it was measured. Stability is claimed only over the evaluated training horizon, and no training horizon length (steps or time) is stated. The source does not say what hardware was used or whether it is NVIDIA Blackwell. Results are for math reasoning on Qwen3 models, so nothing here shows how the method behaves on other tasks.
“TRIAGE modifies the optimization objective while retaining native NVFP4 weight-and activation 4-bit (W4A4) forward execution on both the sampler and learner.”
— TRIAGE paper abstract