CorrGRPO rebalances multi-reward GRPO with Pearson correlations

Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models. It computes advantages by centering and normalizing rewards across rollouts of the same prompt. When there are several rewards, GRPO sums the reward components and divides the total by its within-group standard deviation.
The paper starts from a property of that setup: the variance of the summed reward equals the sum of all pairwise reward covariances. For a fixed centered reward, a larger aggregate covariance produces smaller advantages, and a smaller one produces larger advantages. This lets update magnitudes adapt to how the rewards depend on each other. The authors point to a flaw, though. Correlated rewards with large scales can dominate the normalization and suppress the signals from smaller-scale rewards.
Their fix is Correlation-Normalized GRPO (CorrGRPO). It normalizes the pairwise covariances into Pearson correlation coefficients. According to the authors, the centered total reward stays unchanged, while the influence of differently scaled rewards on the correlation-based normalization is balanced. Advantage magnitudes can still adapt to reward correlations, but the normalization is no longer dominated by large-scale reward components.
For evaluation, the authors compare CorrGRPO with GRPO and other variants on code generation, tool calling and agent security, using models from 0.5B to 8B parameters. They describe these tasks as ones where multiple rewards can improve together or present tradeoffs. They report improvements across all three domains. The code is released at github.com/HKUST-KnowComp/CorrGRPO.
Key facts
- GRPO with several rewards sums them and normalizes by the within-group standard deviation of the total; that variance equals the sum of all pairwise reward covariances.
- The authors argue that correlated rewards with large scales can dominate this normalization and suppress smaller-scale reward signals.
- CorrGRPO normalizes pairwise covariances into Pearson correlation coefficients, keeping the centered total reward unchanged.
- Experiments cover code generation, tool calling and agent security with models from 0.5B to 8B parameters, and the authors report improvements in all three domains.
- Code is available in the HKUST-KnowComp/CorrGRPO repository on GitHub.
Why it matters
GRPO is widely used to train reasoning language models, including with several rewards at once. GRPO's standard handling, summing the rewards and normalizing by the spread of the total, lets rewards with large scales dominate when they are correlated. The paper targets that weak spot with a change to the normalization step only, leaving the centered total reward as it was.
Who it affects
Teams that train language models with GRPO-style reinforcement learning on more than one reward signal. The paper's test domains are code generation, tool calling and agent security, and its models range from 0.5B to 8B parameters.
How to use it
The authors released code at https://github.com/HKUST-KnowComp/CorrGRPO. The method changes how advantages are normalized when several rewards are combined, replacing covariance-based scaling with correlation-based scaling. The summary does not state which model sizes were used for which tasks.
How solid is it
The source is the paper's abstract, and every claim in it is the authors' own. It reports improvements across three domains but gives no benchmark names, no numeric results, no baseline scores and no statistical significance. The other GRPO variants compared against are not named.
Risks and caveats
With no numbers published in the abstract, the size of the gains cannot be judged. No compute cost, training time or overhead of CorrGRPO relative to GRPO is stated. Results come from the authors' own comparisons on three task types, so how well the approach carries over to other reward setups is not shown here.
“correlated rewards with large scales can dominate this normalization and suppress signals from smaller-scale rewards”
— CorrGRPO paper abstract