MeRLa reward shaping cuts RLHF training instability by 41%

Reinforcement Learning from Human Feedback (RLHF) is the standard way to align large language models with human preferences, but the reward models it relies on are static and task-agnostic. That mismatch produces sparse learning signals and suboptimal alignment. A new paper introduces MeRLa (Meta-Learned Reward Shaping), a framework that meta-learns a task-aware shaping function, written as Phi(x, y; phi), across auxiliary tasks before RLHF training even begins. The learned shaping function is combined with the base reward into a composite reward signal. The authors designed the meta-objective around three pieces: task discrimination, entropy regularization, and potential-based conservation, the last of which is meant to keep the shaped reward from changing what the optimal policy actually is, so the extra signal helps training converge without distorting the end goal. The paper also gives theoretical guarantees for this policy invariance, analyzes how sensitive the approach is to representation drift, and works through a specific failure mode where maximizing entropy could create incentives that pull against the intended alignment goal. On the empirical side, the authors tested MeRLa on LLaMA-3-8B across four benchmarks and report consistent improvements over four established RLHF methods: PPO, DPO, GRPO, and DAPO. Headline results include a 90.8% length-controlled win rate on AlpacaEval 2.0, a score of 9.14 on MT-Bench, and 41% less training instability than the baselines. The paper also states that MeRLa keeps delivering these benefits when layered on top of process-based and rubric-based reward enhancements, meaning it is presented as an addition to existing reward setups rather than a replacement that requires discarding them.

Key facts

  • MeRLa meta-learns a task-aware reward-shaping function, Phi(x, y; phi), across auxiliary tasks before RLHF training, addressing the static, task-agnostic reward models used in standard RLHF.
  • Tested on LLaMA-3-8B across four benchmarks, MeRLa shows consistent improvements over PPO, DPO, GRPO, and DAPO.
  • It reaches a 90.8% length-controlled win rate on AlpacaEval 2.0 and a score of 9.14 on MT-Bench.
  • Training instability drops by 41% compared to the baseline methods.
  • The composite reward is built with potential-based conservation to preserve policy optimality, and the gains persist when MeRLa is combined with process-based and rubric-based reward enhancements.

Why it matters

Standard RLHF leans on reward models that are fixed and blind to the specific task at hand, which the authors say leaves the learning signal sparse and the resulting alignment suboptimal. MeRLa targets that root problem directly by having the model meta-learn a shaping function tailored to the task before RLHF training starts, rather than trying to patch the outcome after the fact.

Who it affects

The paper speaks to teams building or fine-tuning RLHF pipelines, particularly anyone currently using PPO, DPO, GRPO, or DAPO to align language models, since the paper reports MeRLa outperforming all four as baselines. The LLaMA-3-8B testbed also makes it directly relevant to groups working at that model scale.

How to use it

MeRLa is a research method, not a shipped product, so there is no price or license to note. As described, it is applied in two stages: first meta-learn the shaping function across auxiliary tasks, then fold the resulting composite reward into RLHF training as the reward signal used to optimize the policy. The paper states the approach keeps working when stacked with process-based and rubric-based reward enhancements, rather than requiring them to be dropped.

How solid is it

The claims rest on experiments across four benchmarks on LLaMA-3-8B, with concrete numbers: a 90.8% length-controlled win rate on AlpacaEval 2.0, a 9.14 score on MT-Bench, and 41% less training instability than PPO, DPO, GRPO, and DAPO. The authors back this with theoretical guarantees for policy invariance and an explicit analysis of representation drift and entropy-driven incentive misalignment, which is more formal grounding than a purely empirical result would carry.

Risks and caveats

The abstract does not name the authors or their institution, and the reported results are confined to one model family, LLaMA-3-8B, and four benchmarks; how MeRLa behaves at other scales or on other model families is not addressed here. As an arXiv posting, the work has not been described as peer-reviewed, and the abstract gives no detail on the added compute or engineering cost of the meta-learning stage itself.

Компания Meta Platforms признана экстремистской организацией, её деятельность на территории РФ запрещена.