Jev-style typed decision models follow labels over definitions, paper finds

Jev-style typed decision models follow labels over definitions, paper finds

A typed decision model answers a fixed question about an input by returning a probability for each of several options that the caller defines. Each option has a short label and a written definition, and the definition is where a developer states the rule the model should apply. Jev introduced this interface for routing, moderation and triage, and open implementations followed. The paper studies those open implementations, because their weights can be inspected and patched, and asks whether the probability follows the definitions or the labels. The authors call a preference for the label option-label bias.

Their answer is that the labels win. The finding holds across four open-weight typed decision models, three ways of reading an answer from a Qwen2.5 backbone, eleven classification tasks, and PolicyBench, a synthetic routing suite the authors introduce in which the rule appears only in the definitions.

The numbers show it plainly. Deleting every definition leaves accuracy unchanged (laya-td: 0.8559 against 0.8487), even though those definitions support 0.7971 on their own. Renaming the options to A and B raises accuracy by +0.1511 [+0.1377, +0.1646].

One system, von, is unaffected, and the two code bases differ in one expression. laya writes each option as "{label}: {definition}", while von writes only the definition. Changing that expression in both directions, with no weight changed, makes all three laya checkpoints exactly invariant (+0.0000 [+0.0000, +0.0000]) and creates the effect in von, whose accuracy falls from 0.8511 to 0.2281 when a label contradicts its definition.

Earlier work blamed the constrained decision head these models use in place of a text decoder. The authors say their results locate the problem in the prompt rendering instead. They also give a two-call test that tells a practitioner which case applies to their model, and they measure what four mitigations are worth. They add that the same operation occurs whenever a language model is used as a classifier by scoring label strings.

Key facts

  • Across four open-weight typed decision models, three readings from a Qwen2.5 backbone, eleven tasks and the new PolicyBench suite, the answer is mostly determined by the labels, not the definitions.
  • Deleting every definition leaves accuracy unchanged (laya-td: 0.8559 against 0.8487), although the definitions alone support 0.7971; renaming options to A and B raises accuracy by +0.1511 [+0.1377, +0.1646].
  • The two code bases differ in one expression: laya writes "{label}: {definition}" for each option, von writes only the definition, and von is unaffected.
  • Changing that expression in both directions, with no weight changed, makes all three laya checkpoints exactly invariant and creates the effect in von, whose accuracy falls from 0.8511 to 0.2281 when a label contradicts its definition.
  • The authors locate the failure in prompt rendering rather than the constrained decision head blamed by earlier work, and offer a two-call test plus measurements of four mitigations.

Why it matters

Typed decision models are meant for routing, moderation and triage, where the written definition of each option is the rule a developer wants applied. If the probabilities follow the label strings instead, the rule the developer wrote may do little or nothing. The paper shows this concretely: deleting every definition leaves accuracy unchanged, even though the definitions on their own support 0.7971. It also revises the diagnosis. Earlier work pointed at the constrained decision head used in place of a text decoder; the authors' results point at the prompt rendering, and changing one expression, with no weight changed, was enough to remove the effect in laya and create it in von.

Who it affects

Developers who build routing, moderation or triage on open implementations of Jev's typed decision model interface are the direct audience, since the study covers four open-weight models. The authors also say the same operation occurs whenever a language model is used as a classifier by scoring label strings, so anyone classifying that way has reason to look at how their options are written.

How to use it

The practical advice in the source is to find out which case applies to your model. The authors give a two-call test for that purpose, and they measure what four mitigations are worth. The results of those mitigations and the details of the test are not given in the source text. The clearest lever the paper shows is the rendering expression: laya writes each option as "{label}: {definition}" and is affected, while von writes only the definition and is not. Renaming options to A and B raised accuracy by +0.1511 in the authors' measurements.

How solid is it

The evidence is quantitative and spans four open-weight models, three ways of reading an answer from a Qwen2.5 backbone, eleven classification tasks and the authors' own PolicyBench. Effects come with confidence intervals, and the intervention is clean: the rendering expression was changed in both directions with no weight changed, and the result flipped between the two code bases. The source is the paper's abstract-length summary, so the full methods are not visible here. No authors or institutions are named in it. The authors work only with open implementations whose weights can be inspected and patched.

Risks and caveats

The source does not say these findings apply to closed models, only to open implementations whose weights can be inspected and patched. The four mitigations' results and the content of the two-call test are not given. The source also does not say which specific Qwen2.5 size is used, or name the four models beyond laya-td and von. The von figure of 0.8511 is given only as the starting value in the contradicting-label condition, not as a baseline for a particular task set. The pair 0.8559 and 0.8487 is given as the comparison for deleting every definition, without saying which value belongs to which condition.

“Earlier work attributed this failure to the constrained decision head these models use in place of a text decoder; our results locate it in the prompt rendering.”

— Paper abstract, Hugging Face Papers