R-Quest keeps self-evolving reasoning models from collapsing by filtering bad questions

R-Quest keeps self-evolving reasoning models from collapsing by filtering bad questions

Self-evolving reasoning models improve by learning from questions they generate themselves, but the paper opens with a known problem: repeated self-training can lead to performance collapse. The authors set out to explain why performance deteriorates over successive rounds and how to keep self-evolution going.

Their analysis finds two recurring quality problems in the self-generated questions. The first is invalid questions. These become more prevalent across rounds, and answer-consistency filtering, a common way to clean the data, further increases their proportion in the training data. The second is repetition. Existing controls on question diversity rely on lexical similarity, so they can miss mathematically equivalent questions that are worded differently. The result is a collapse in question diversity in later training rounds.

Building on these findings, the authors introduce R-Quest, which uses question validity and novelty feedback to guide self-evolution. For validity, they first train the solver to recognize and reject invalid questions, then use its judgments in two ways: to guide the questioner's rewards and to filter the solver's training data. For novelty, a frozen base model compares sampled pairs of questions and provides feedback, which is meant to avoid question repetition.

On results, the paper states that R-Quest consistently achieves the highest average performance on 12 benchmarks covering mathematical reasoning, general-domain reasoning and code generation, across two model families. It also maintains stable performance gains over ten rounds of self-evolution, peaking in the final round and outperforming R-Zero by 17.32 points.

Key facts

  • Repeated self-training in self-evolving reasoning models can lead to performance collapse; the paper investigates why.
  • Two problems are identified in self-generated questions: invalid questions that grow across rounds (and rise further under answer-consistency filtering), and mathematically equivalent repeats that lexical-similarity diversity controls miss.
  • R-Quest trains the solver to reject invalid questions, uses its judgments for questioner rewards and training-data filtering, and uses a frozen base model to compare question pairs for novelty feedback.
  • It reports the highest average performance on 12 benchmarks (mathematical reasoning, general-domain reasoning, code generation) across two model families.
  • Gains stay stable over ten self-evolution rounds, peaking in the final round, where R-Quest outperforms R-Zero by 17.32 points.

Why it matters

Self-evolution, where a model writes its own training questions and learns from them, is attractive because it needs little outside data. The catch is that it can fall apart after several rounds. This paper offers a specific diagnosis rather than a vague warning: the question generator drifts toward invalid questions and toward rewordings of the same problem. It also points out that a standard filter, answer-consistency filtering, makes the invalid share of the training data larger. Knowing where the decay comes from is what makes a targeted fix possible.

Who it affects

Mainly researchers building self-play or self-evolving training loops for reasoning models, where a questioner generates problems and a solver learns from them. The reported evaluation spans mathematical reasoning, general-domain reasoning and code generation, so the claim reaches beyond maths alone. R-Zero is the baseline the paper compares against.

How to use it

The recipe is described in the abstract. Train the solver to recognize and reject invalid questions. Use those judgments to shape the questioner's rewards and to filter the solver's training data. Separately, have a frozen base model compare sampled question pairs and give novelty feedback, so that semantically equivalent questions are not counted as different just because their wording differs. No code, dataset or model release is mentioned in the source.

How solid is it

The claims come from the paper's own abstract, so they are the authors' self-reported results. The headline numbers are 12 benchmarks, two model families, ten rounds and a 17.32-point margin over R-Zero. The metric behind that margin is not stated, nor whether it is absolute accuracy or an average across benchmarks. The 12 benchmarks are not named, and no model names or sizes are given for the two families, so the strength of the result cannot be judged from this text alone.

Risks and caveats

The 17.32-point figure is given for the final round of ten, and the source does not say which metric it refers to, so it should not be read as a general accuracy gain. The method adds machinery: a solver trained to judge validity and a frozen base model for pairwise novelty comparison. The source gives no cost figures for this. The results cover two model families over ten rounds, and the source does not say whether the gains hold beyond that.

“repeated self-training can lead to performance collapse”

— Abstract of "Questioning the Questions: Sustaining Self-Evolution in Reasoning Models"