Language models get derailed by changes users later rejected, paper finds

Language models get derailed by changes users later rejected, paper finds

A paper listed on Hugging Face Papers (2610.06496) starts from a simple expectation. When a user is working with a language model on a multi-turn task, proposes a change and then rejects it, the model should carry on as if nothing had changed.

The authors report that this does not hold. In their words, merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged.

To study the problem in a controlled way, they introduce Intent-Eval, a benchmark spanning tool actions, code, databases and mathematics. Across these diverse tasks, they find that models are vulnerable to two kinds of conversational events: proposals the user rejected, and requirements the user superseded with new ones.

The authors describe the pattern as mentioned-as-in-effect confusion. Conversational content is treated as an active requirement even after it has been rejected or replaced. They also report that accuracy degradation can deepen or persist as the interaction continues, so the damage does not necessarily fade over later turns. Their takeaway is that models need to distinguish what has been mentioned from what remains in effect.

Building on that, the paper proposes Intent-OPSD, a decision-conditioned on-policy self-distillation framework. The Teacher and the Student are initialized from the same model. The frozen Teacher supplies active-intent supervision from the complete task that matches the user's decision, and the Student is trained on the full dialogue to follow the active requirements that reflect the user's intent.

Key facts

  • The authors report that merely mentioning a rejected change can derail a model's task execution, even when the user's final intent is unchanged.
  • Intent-Eval is a controlled benchmark spanning tool actions, code, databases and mathematics.
  • Models are vulnerable to both rejected proposals and superseded requirements, which the authors call mentioned-as-in-effect confusion.
  • Accuracy degradation can deepen or persist as the interaction continues.
  • Intent-OPSD is an on-policy self-distillation framework in which a frozen Teacher and a Student start from the same model, and the Student learns to follow active requirements.

Why it matters

Real conversations with a model are full of ideas that get floated and dropped. The paper's point is that a model should treat a rejected change as if it had never been requested, and that current models reportedly do not. The failure is described as surprising because the user's final intent is the same as before; only the mention of the rejected change differs. The authors name the underlying problem as mentioned-as-in-effect confusion: text that appeared in the dialogue is treated as an active requirement even after it has been rejected or replaced. Their argument is that models need to separate what has been mentioned from what remains in effect.

Who it affects

The failure is reported across tool actions, code, databases and mathematics, so it concerns anyone who steers a model through a multi-turn task and changes direction along the way. That includes people who propose an option, then decide against it, and people who replace an earlier requirement with a new one. It is also relevant to those who build or evaluate assistants that run such tasks, since Intent-Eval is framed as a way to study behavior under evolving user intent.

How to use it

The source text mentions no release of code, data or benchmark, so there is nothing described here to download or run. What the paper offers is a way of thinking about the problem. When testing a model or an assistant on a multi-turn task, include dialogues in which the user proposes a change and then rejects it, and dialogues in which a requirement is replaced, and check whether the final answer still follows the user's active intent. The proposed Intent-OPSD recipe, as described, trains a Student on the full dialogue with supervision from a frozen Teacher that sees the complete task matching the user's decision.

How solid is it

The text available is the paper's abstract-level summary, and every claim in it is the authors' own. It gives no numerical results: no accuracy figures, no sizes of degradation and no benchmark sizes. It does not name which models were evaluated. It does not report any result for Intent-OPSD, so it does not say whether the method fixes the failure or by how much. No authors or institutions are named in the text.

Risks and caveats

The headline finding and the proposed fix are both stated without supporting numbers, so the size of the problem and the benefit of Intent-OPSD cannot be judged from this text. The claim that degradation can deepen or persist as interaction continues is phrased as a possibility, not a fixed rule. Because the models tested are not named, it is unclear which systems the failure applies to. Treat the paper as a described problem and a proposed approach, not as a demonstrated solution.

“merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged”

— Abstract of the paper, Hugging Face Papers 2610.06496