SanSi turns a looped language model into a decision model with 72.0% accuracy

SanSi turns a looped language model into a decision model with 72.0% accuracy

The paper starts from typed decision models: models that answer a declared question without generating text. A decision head returns a probability for each declared option in a single forward pass. The authors describe that single pass as fast, intuitive System 1 thinking, and they study what lies between one pass and fully generated reasoning.

Their answer is looping, in which the same layers are applied recursively several times before one typed readout. Each loop lets the model revise its hidden state before it commits to an answer, and no token is generated along the way. The authors call this System 1.5 thinking.

The method is called SanSi. It turns a pre-trained looped language model into a typed decision model. Option probabilities are read after every loop, and every loop is trained with a proper scoring rule. As a result, one model serves every loop budget from one loop to eight, and it does so in a single training run.

On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy. That is 13.5 points above a non-looped model of the same shape trained with the same recipe, and 5.3 points above a newer non-looped model of its size. It sits 1.8 points below a model with three times the parameters. All three gaps are percentage points.

Two further results are reported. On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails. And when SanSi is used as the judge for policy optimization with reinforcement learning, without gold answers, it raises the generator's F1 by 7.7 points.

Key facts

  • SanSi converts a pre-trained looped language model into a typed decision model that returns a probability for each declared option without generating text.
  • Option probabilities are read after every loop and each loop is trained with a proper scoring rule, so one model covers loop budgets from one to eight in a single run.
  • On 10,027 test decisions from 59 sources, SanSi reaches 72.0% accuracy: 13.5 points above a non-looped model of the same shape and recipe, 5.3 points above a newer non-looped model of its size, and 1.8 points below a model with three times the parameters.
  • On two depth-controlled tasks, loops extend the solvable depth beyond the depths seen in training, where the larger single-pass model fails.
  • Used as the judge for reinforcement-learning policy optimization without gold answers, SanSi raises the generator's F1 by 7.7 points.

Why it matters

Decision models that answer in one forward pass are fast but have no way to think longer, while generated reasoning costs tokens. SanSi targets the space between: extra computation happens inside the network by looping the same layers, and nothing is written out. The authors call this System 1.5 thinking. The headline claim is that a looped model beats a non-looped model of the same shape by 13.5 points and comes within 1.8 points of a model with three times the parameters.

Who it affects

The work is aimed at people building decision-style components, where a model must choose among declared options and return probabilities. It also touches those who train generators with reinforcement learning, since the abstract reports SanSi serving as a judge that needs no gold answers and lifting the generator's F1 by 7.7 points.

How to use it

The abstract describes a recipe rather than a product: start from a pre-trained looped language model, add a typed decision head, read the option probabilities after every loop, and train each loop with a proper scoring rule. The loop budget can then be chosen anywhere from one to eight with a single model. No release of code, weights or data is mentioned.

How solid is it

This is a preprint abstract, and the figures below are the authors' own. The evaluation covers 10,027 test decisions from 59 sources, with SanSi at 72.0% accuracy. The comparisons are against a same-shape non-looped model trained with the same recipe, which is the cleanest test of looping itself, plus a newer non-looped model of its size and a model with three times the parameters. The abstract names no authors or institutions.

Risks and caveats

The abstract does not name the baseline models or give parameter counts, and it does not say which base looped model SanSi is built on. The nature of the 59 sources and the decision tasks is not described, and the two depth-controlled tasks are not named. The 72.0% figure is not tied to a specific number of loops, and accuracy at individual loop counts is not reported. The 7.7-point F1 gain has no stated baseline value or task. No speed, latency or compute-cost figures are given for loops versus generated reasoning.

“Each loop lets the model revise its hidden state before it commits to an answer, without generating a token; we call this System 1.5 thinking.”

— SanSi paper abstract