Nereus adapts RL post-training plans mid-run, beating Verl and OpenRLHF

Nereus adapts RL post-training plans mid-run, beating Verl and OpenRLHF

Reinforcement learning (RL) post-training for large language models coordinates several models across generation, inference and training stages on a GPU cluster. The paper behind this story starts from a simple observation: conditions change during a run. Resource availability, sequence length, memory pressure and stage bottlenecks can all shift, so an execution plan that was a good fit at the start can become slow, or even infeasible, later on.

Changing the plan mid-run is hard when the job's models share GPUs. The authors name three difficulties: deciding whether a new plan is worth the cost of the transition, reusing the job's distributed state, and coordinating GPU transfers across models and stages.

Their answer is Nereus, described as a cost-aware runtime that adapts RL post-training jobs into efficient execution plans. It has a low-overhead controller that selects a memory-feasible global plan and admits a transition only using a cost model calibrated against the running job. To estimate and carry out a transition, Nereus represents the distributed state of each replica of a model-stage (one model in one stage) as an Elastic Model Unit. It then uses a global transition graph to order the transformations and GPU transfers of those units.

The abstract reports three results. In a trace built from real data, online TP/PP adaptation reduces average step latency by 27.7% relative to the initial fixed TP/PP layout with DP scaling. In a 1,000-step run that reaches 1,024 GPUs, six transitions consume 0.079% of total run time. And across diverse clusters, Nereus improves end-to-end 8B PPO throughput by 2.14 to 7.27 times over OpenRLHF and by 1.10 to 1.47 times over Verl.

Key facts

  • Nereus is a cost-aware runtime that changes the execution plan of an RL post-training job for LLMs while the job is running.
  • In a trace built from real data, online TP/PP adaptation cut average step latency by 27.7% versus the initial fixed TP/PP layout with DP scaling.
  • In a 1,000-step run reaching 1,024 GPUs, six transitions took 0.079% of total run time.
  • End-to-end 8B PPO throughput rose by 2.14 to 7.27 times over OpenRLHF and by 1.10 to 1.47 times over Verl across diverse clusters.
  • Each model-stage replica is modelled as an Elastic Model Unit, and a global transition graph orders the state transformations and GPU transfers.

Why it matters

RL post-training runs on clusters where the best split of work between generation, inference and training does not stay fixed. Per the abstract, a plan that starts well can become slow or infeasible as resources, sequence length, memory pressure and stage bottlenecks change. Nereus treats the plan as something to revise during the run, and only when a cost model says the switch pays for itself. The reported overhead is small: six transitions used 0.079% of the time in a 1,000-step run.

Who it affects

Mainly teams that run RL post-training of LLMs on shared GPU clusters, and the people who build the training frameworks they use. The paper compares against two existing systems, OpenRLHF and Verl, so users of those frameworks are the most direct audience.

How to use it

The abstract describes a research system, not a product. No code release, license or availability is mentioned. For practitioners, the useful part is the idea: check whether a running job's parallelism layout still fits, and change it only when the estimated gain outweighs the transition cost.

How solid is it

The evidence is the numbers in the abstract, which are the authors' own. The throughput comparison covers 8B PPO only, against OpenRLHF and Verl, across diverse clusters. The 27.7% latency figure comes from a trace built from real data, not a live run. The abstract does not say which cluster gives the minimum or maximum of the speedup ranges, and it names no authors or institutions.

Risks and caveats

The gap to Verl (1.10 to 1.47 times) is far narrower than the gap to OpenRLHF (2.14 to 7.27 times), so the size of the benefit depends heavily on the baseline. No hardware type or cluster specifications are given beyond the 1,024-GPU figure, and no other model sizes or algorithms are stated. The 0.079% overhead was measured in one 1,000-step run with six transitions, so it is a single data point.