KLPO proposes critic-free RL for LLM agents with one rollout per prompt

The paper starts from a problem in asynchronous reinforcement learning for LLM agents. One policy is trained on trajectories produced by another: the rollouts come from stale checkpoints, and the inference engine's probabilities differ from the trainer's even when the parameters are identical. The authors say the standard fixes both have a cost. Clipping importance ratios biases the update. Sampling a group of responses per prompt, as GRPO does, is expensive when episodes are long.
The authors propose KL-Regularized Policy Optimization (KLPO), a framework that anchors the KL regularizer at the sampler. With that anchor, the regularized improvement step has a closed-form Gibbs solution. KLPO then fits the log-ratio optimality condition of that solution by least squares, using the sampler's own trajectories. Because the sampler probability enters only through a log-ratio, no importance weights are needed.
The intractable part of the objective is the log-partition function. The paper profiles out the regression intercept, which replaces the log-partition function with the signal's sampler mean plus a sampler-to-trainer KL divergence.
For token-level policy mirror descent targets, the authors show the resulting gradient can be computed from terminal returns without a critic. They do this either with sampler-centered scores or with a single trajectory residual, and the result holds even under stochastic tool outputs.
Three further results are listed. The authors prove that independent Monte Carlo estimates of the KL term keep these gradients unbiased. They derive the exact KL gap of cheaper top-K and binary approximations. And they show that four existing methods, SPPO, GPO, REBEL and BPO, arise as special cases of KLPO.
The end product, in the authors' words, is a critic-free update that uses one rollout per prompt and requires neither a learned normalizer nor a group of responses.
Key facts
- KLPO targets asynchronous RL for LLM agents, where rollouts come from stale checkpoints and inference-engine probabilities differ from the trainer's even at identical parameters.
- It anchors the KL regularizer at the sampler, so the sampler probability enters through a log-ratio and no importance weights are needed.
- The update is critic-free and uses one rollout per prompt, in contrast to GRPO, which samples a group of responses per prompt.
- The authors prove that independent Monte Carlo estimates of the KL term keep the gradients unbiased, and derive the exact KL gap of cheaper top-K and binary approximations.
- SPPO, GPO, REBEL and BPO are shown to arise as special cases of KLPO.
Why it matters
Asynchronous training is a practical reality for LLM agents: the policy that generates trajectories lags behind the one being trained, and the inference engine can disagree with the trainer even at identical parameters. The authors argue that the usual answers each carry a price. Clipped importance ratios bias the update, and GRPO-style group sampling gets expensive when episodes are long. KLPO is offered as a way out of both: no importance weights and no group of responses, just one rollout per prompt and no critic. Showing that SPPO, GPO, REBEL and BPO are special cases also positions it as a unifying view of several earlier methods.
Who it affects
The paper speaks to people training LLM agents with asynchronous reinforcement learning, especially where episodes are long and sampling a group of responses per prompt is costly. It is also relevant to researchers working on KL-regularized and policy-mirror-descent style objectives, since the paper connects four existing methods to its framework.
How to use it
No code release or implementation details are mentioned in the source, so there is nothing to run yet. Conceptually, the recipe in the abstract is: sample one rollout per prompt from the sampler, anchor the KL regularizer at that sampler, and fit the log-ratio optimality condition by least squares on those same trajectories. The gradient can then be computed from terminal returns, using either sampler-centered scores or a single trajectory residual, with no learned critic or normalizer. The paper also covers cheaper top-K and binary approximations of the KL term and derives their exact KL gap, which is the part to check before choosing a cheaper estimator.
How solid is it
The source is the paper's abstract, so what is available is the authors' own account of their theory: a closed-form Gibbs solution, an unbiasedness proof for independent Monte Carlo estimates of the KL term, an exact KL gap for the cheaper approximations, and the special-case relationships to SPPO, GPO, REBEL and BPO. No experimental results, benchmarks, model sizes or numerical improvements are reported in the abstract. The abstract does not say whether KLPO has been tested on real LLM agent tasks.
Risks and caveats
Every claim here is the authors' own, and the abstract gives no comparison of KLPO's wall-clock cost, accuracy or stability against GRPO or other baselines. Whether one rollout per prompt matches group-based methods in practice is therefore not established by the source. The cheaper top-K and binary approximations come with a KL gap, which the authors derive exactly, so using them trades some accuracy of the KL term for lower cost. No authors or institutions are named in the source.
“The result is a critic-free update that uses one rollout per prompt and requires neither a learned normalizer nor a group of responses.”
— KLPO paper abstract