ScholarCatalyst benchmark tests whether AI can find papers that inspire research

A new benchmark called ScholarCatalyst targets a skill the authors say scientists still have over AI systems: sensing which earlier idea, buried in an ever-growing archive of research, a new problem needs. Even as AI starts to make progress on open problems, the authors write, scientists remain far ahead at this.
To study it, the authors turned to people who know firsthand which earlier work helped their projects: the researchers themselves. Papers serve as pointers to the ideas inside them. Using an automated pipeline that makes author annotation scalable, they had 184 lead authors of 207 recent computer science papers label which candidate papers did or could have advanced their project. Each label comes with a detailed rationale.
The benchmark task is a retrieval problem with author-provided judgments. Given only an initial research question, a system must retrieve those papers from just the literature that was available when the project began.
The results are modest. Agentic search does no better than embedding retrieval: 0.42 Recall@20 against 0.48, even though the agent calls that same retriever as a tool. Even an agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20.
The authors conclude that the results highlight a need for new training recipes that equip models with expert intuition for searching broad corpora. They describe ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.
Key facts
- ScholarCatalyst is built from labels by 184 lead authors of 207 recent computer science papers, who marked which candidate papers did or could have advanced their project, each with a detailed rationale.
- The task: given an initial research question, retrieve the author-judged papers using only the literature available when the project began.
- Agentic search scored 0.42 Recall@20 against 0.48 for embedding retrieval, even though the agent calls that same retriever as a tool.
- An agent built on Claude Fable 5.1, which may have seen the completed papers during training, reaches only 0.51 R@20.
- The authors call for new training recipes that give models expert intuition for searching broad corpora.
Why it matters
Finding the right earlier idea for a new problem is a core part of research, and the authors argue AI systems are still far behind scientists at it. Most retrieval tests ask for papers on a topic. This one asks for papers that actually advanced a project, as judged by the people who did the work. The headline result is that giving a model tools does not help much: the agent that calls an embedding retriever scored 0.42 Recall@20, below the 0.48 of the retriever alone.
Who it affects
Researchers building scientific agents and literature-search tools are the direct audience, since the benchmark measures the skill those tools need. Scientists who hope for an assistant that can take a half-formed idea and point to the prior research it needs are affected too, because the results suggest current systems fall well short of that.
How to use it
ScholarCatalyst defines a concrete test: give a system an initial research question and have it retrieve the author-judged papers from only the literature available when the project began, then score it with Recall@20. Teams working on agentic search or training could use that setup to compare approaches. The source says: "No release, dataset link or code availability is stated."
How solid is it
The labels come from first-hand knowledge: 184 lead authors of 207 recent computer science papers marked which candidates did or could have advanced their projects and gave a detailed rationale for each. The scores are reported as stated by the authors, and the Claude Fable 5.1 figure comes with the authors' own note that the model may have seen the completed papers during training. This account rests on the paper's abstract-level description. The source does not name the embedding retriever or the other agents or models evaluated beyond the Claude Fable 5.1 agent.
Risks and caveats
The claim that scientists remain far ahead of AI at this skill is the authors' framing; the figures given are for AI systems only. The 0.51 R@20 of the Claude Fable 5.1 agent is compared with 0.48 for embedding retrieval, and the source does not state whether that agent used the same retriever. The source does not say how many candidate papers per project were labelled or how many papers are in the corpus, nor which fields of computer science the papers come from. The benchmark covers recent computer science papers, so results may not carry over to other fields.
“We envision ScholarCatalyst as a step toward scientific agents that can take a half-formed idea and point to the prior research it needs.”
— ScholarCatalyst authors