GAGAR paper: agentic grader ranks passing code, reweights RL rewards

GAGAR paper: agentic grader ranks passing code, reweights RL rewards

A paper on Hugging Face introduces GAGAR, a framework for quality-aware credit redistribution in reinforcement learning (RL) for code agents. It targets a weakness in a common setup. RL for code agents often uses executable tests to give binary rewards: pass or fail. With such rewards, Group Relative Policy Optimization (GRPO) assigns identical advantages to every test-passing trajectory within a rollout group. That ignores differences in implementation quality and in how closely each solution follows the task requirements. The authors say this leaves the policy without a learning signal that favors clean, targeted implementations over those with unnecessary or out-of-scope changes.

GAGAR builds on dynamic sampling, which keeps groups that contain both passing and failing trajectories. It then places all trajectories of a group in a shared workspace. There, an agentic grader trained with supervised fine-tuning (SFT) inspects them jointly and ranks the test-passing candidates. Based on that ranking, lower-ranked trajectories are downweighted, and the advantages of all test-passing trajectories are proportionally rescaled so they add up to their original sum. According to the authors, this sum-preserving redistribution keeps the relative weights set by the quality-based downweighting while shifting credit toward higher-quality implementations.

The evaluation is described as industrial scale. It uses pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters). In controlled code-only experiments with Flash, the authors report improved code agent performance, reduced growth in trajectory length, and more stable training. They also apply GAGAR in large-scale mixed-task RL with both Flash and Pro. They conclude that their results support combining test-based verification with groupwise agentic grading to improve the quality and stability of code agent RL.

Key facts

  • GRPO with binary test rewards gives identical advantages to all test-passing trajectories in a group, so it cannot tell clean patches from ones with unnecessary or out-of-scope changes.
  • GAGAR puts a group's trajectories in a shared workspace, where an SFT-trained agentic grader jointly inspects them and ranks the test-passing ones.
  • Lower-ranked passing trajectories are downweighted, and all passing advantages are rescaled to restore their original sum.
  • Evaluated on pre-RL SFT checkpoints of MiMo-V2.6-Flash (310B total parameters) and MiMo-V2.6-Pro (1.02T total parameters).
  • Controlled code-only Flash experiments report better code agent performance, slower trajectory-length growth and more stable training.

Why it matters

Test-based rewards are a standard way to train code agents, but a pass/fail signal treats a tidy, minimal fix and a sprawling one with out-of-scope edits as equal. GAGAR tries to add a quality signal on top of tests without dropping the tests as the ground truth. The sum-preserving rescaling means the total credit given to passing trajectories stays the same; only its distribution among them changes.

Who it affects

Teams that train code agents with RL and GRPO-style methods, especially at large scale. The paper's own evaluation uses MiMo-V2.6-Flash and MiMo-V2.6-Pro checkpoints.

How to use it

The abstract describes the recipe. Keep groups with both passing and failing trajectories through dynamic sampling. Have an SFT-trained agentic grader inspect the whole group in a shared workspace and rank the passing candidates. Downweight the lower-ranked ones, then rescale all passing advantages so their sum matches the original. The text does not mention released code or models.

How solid is it

The claims come from the paper's abstract and are the authors' own. Improvements are stated only qualitatively: better performance, reduced trajectory-length growth, more stable training. The controlled experiments are code-only and use Flash. The mixed-task runs with Flash and Pro are described only as GAGAR being applied. No benchmark names or numeric results are given.

Risks and caveats

The method depends on an agentic grader, and the text does not state its size, base model or training data. Results for the 1.02T Pro model are not described beyond the fact that GAGAR was applied. Without numbers, the size of any gain cannot be judged from this text.