QwenGyre framework speeds up RL training for xlong-horizon LLM agents

QwenGyre framework speeds up RL training for xlong-horizon LLM agents

A new paper presents QwenGyre, an end-to-end framework for online reinforcement learning (RL) on what the authors call extreme-long, or xlong, horizon tasks for LLM agents. The authors say agents increasingly take on such tasks, where a single execution can span hours, hundreds of model-environment interactions, and nearly 1M tokens per rollout.

The paper names two fundamental problems when online RL is applied to executions of this kind. First, severe execution variance and prolonged rollout delays cause massive GPU idling. Second, complex non-linear branching generates massive trajectory redundancy, which the authors say cripples training efficiency.

QwenGyre answers each problem with its own mechanism. For the idle GPUs, it elastically reallocates GPUs between rollout and training without interrupting live executions. For the redundancy, its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths, which the authors say bounds training costs.

The paper gives two sets of results. Scaled to the authors' flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench, from 52.5% to 58.5%, in 48 steps. Across evaluations on diverse domains of training datasets, QwenGyre delivers up to 1.85× and 1.78× speedups over the Colocate and Async baselines, respectively.

Key facts

  • QwenGyre is an end-to-end framework for online RL on xlong-horizon LLM agent tasks, which can span hours, hundreds of model-environment interactions and nearly 1M tokens per rollout.
  • It targets two problems: GPU idling caused by execution variance and rollout delays, and trajectory redundancy caused by non-linear branching.
  • It elastically reallocates GPUs between rollout and training without interrupting live executions, and its trajectory processor reconstructs branching histories, scores partial progress and deduplicates redundant paths.
  • On the authors' flagship Qwen~3.8 2.4T with 700K tokens per rollout, NL2RepoBench rises from 52.5% to 58.5% (a 6.0% absolute gain) in 48 steps.
  • Speedups reach up to 1.85× over Colocate and up to 1.78× over Async across diverse training datasets.

Why it matters

Agent tasks that run for hours and use close to a million tokens per rollout are hard to train with online RL, because a few slow executions can leave GPUs sitting idle and branching histories inflate the data to be trained on. QwenGyre attacks both problems in one framework: it moves GPUs between rollout and training on the fly, and it trims redundant branches. The reported result is faster training against two baselines, Colocate and Async, plus a measured benchmark gain on the authors' largest model.

Who it affects

The paper is aimed at teams that train LLM agents with online RL on very long tasks, where rollout time and trajectory volume drive the cost. The benchmark result is reported on the authors' own flagship model, Qwen~3.8 2.4T, so the direct evidence concerns that model.

How to use it

The abstract describes the design but gives no usage steps. No release, availability, code or open-source status of QwenGyre or Qwen~3.8 2.4T is stated. In practice, the paper is a description of an approach: elastic GPU reallocation between rollout and training, plus a trajectory processor that reconstructs branching histories, scores partial progress and deduplicates paths.

How solid is it

These are the authors' own reported numbers from a paper abstract. The 6.0% absolute gain on NL2RepoBench (52.5% to 58.5% in 48 steps) is reported only for the flagship model at 700K tokens per rollout. The speedups of 1.85× over Colocate and 1.78× over Async are 'up to' maxima across evaluations on diverse training datasets; no typical or average speedup is given.

Risks and caveats

The source does not say which datasets or domains were used in the speedup evaluations, nor what Colocate and Async are beyond being baselines. It does not say whether the baseline comparison was run on the same model and hardware as the NL2RepoBench result. No wall-clock time, GPU count or cost figures are given. The NL2RepoBench gain is reported only for the flagship model; no results for other models are given.