Jev beats Laya on 9 of 11 agent decision points in System-1 model audit
Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries a prompt injection. System-1 decision models answer such questions in a single forward pass and return class probabilities. They promise large cost and latency savings compared with calling an LLM for each decision. This arXiv paper tests that promise.\n\nThe authors compare an open-weight System-1 model, Laya, with a hosted one, Jev. The test covers 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants. The setup uses byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks.\n\nJev is significantly more accurate on 9 of the 11 decision points, with advantages from +10.8 to +46.0 percentage points. Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating. Laya is also fragile: it changes 30% of its answers when the option order is reversed, and it degrades sharply when there are many or similar candidates. The paper cites 31% for Laya at 50 nearest-neighbour tools, against 98% for Jev on items with a unique correct tool.\n\nThe second half of the paper is a self-audit. The authors say three analysis errors and one design confound distorted their headline deployment claims. First, an omitted pre-screen cost: they had reported a 23.9% saving, and the actual figure is 4.3%. Second, gate accuracy had been reported as end-to-end quality (58% vs. 98%). Third, thresholds were set in-sample: the target was a 5% miss rate, but held-out misses reached up to 17%. Fourth, a "channel effect" on injection false positives vanishes once the content is channel-native. Two other suspected confounds did not change the conclusions.\n\nAll cases, raw outputs and analysis code are available in a public GitHub repository at github.com/David-DL-Space/sys1-eval.
Key facts
- The paper compares open-weight Laya and hosted Jev on 11 agent decision points built from 18 public sources, with 7,283 base cases and 6,640 robustness variants.
- Jev is significantly more accurate on 9 of 11 decision points, by +10.8 to +46.0 percentage points; neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating.
- Laya changes 30% of its answers when option order is reversed and degrades sharply with many or similar candidates.
- The authors audit their own earlier claims: a reported 23.9% saving was actually 4.3% once pre-screen cost was included, and in-sample thresholds aiming at 5% misses hit up to 17% on held-out data.
- All cases, raw outputs and analysis code are public on GitHub.
Why it matters
System-1 decision models are pitched as a cheap, fast replacement for LLM calls on the small routing, tool-choice, relevance and injection checks inside agent harnesses. This paper puts that pitch under a controlled test and finds big gaps between the two models and some outright failures: neither beats chance on zero-shot model routing. It also shows how easily savings claims go wrong. A 23.9% saving became 4.3% when an omitted pre-screen cost was counted.
Who it affects
Teams building agent harnesses who are considering a single-forward-pass model for decisions such as model routing, tool selection, RAG relevance gating or injection detection. It also matters to anyone who publishes or relies on cost and latency savings numbers for such components, since the audit points to specific places where those numbers drifted.
How to use it
The authors publish all cases, raw outputs and analysis code at github.com/David-DL-Space/sys1-eval, so the evaluation can be rerun or extended to other decision points. The method itself is a usable checklist: byte-identical inputs across models, paired tests, reversed option order, large or near-duplicate candidate sets, and reproducibility checks across hardware and days. The paper also warns about three reporting traps: count pre-screen costs, report end-to-end quality rather than gate accuracy, and set thresholds on held-out data rather than in-sample.
How solid is it
The design is careful: paired tests on 7,283 base cases and 6,640 robustness variants, with cross-hardware and cross-day reproducibility checks, and public data and code. The authors also report their own mistakes and say two other suspected confounds did not change the conclusions. These findings come from the authors' own abstract and have not been independently checked here.
Risks and caveats
The results cover two models, Laya and Jev, on 11 decision points, so they are not a verdict on System-1 models in general. The 31% figure for Laya at 50 nearest-neighbour tools is given in the abstract without an explicit metric, so it should not be read as a precise accuracy comparison. The corrected figures (23.9% vs. 4.3%, 58% vs. 98%) are the authors' own, and the abstract does not say where the earlier claims were published. The audit also shows how much the results depend on thresholds, option order and content type.
“Neither model beats chance on zero-shot model routing, and they tie on RAG relevance gating.”
— Paper abstract, arXiv 2610.02267