Station agent environment rediscovers 62.7% of findings from three ICLR papers

A new paper asks whether AI agents can do open-ended scientific discovery, not just optimise against a well-defined metric. The authors note that recent AI systems have made rapid progress in discovery when given clear metrics, but whether they can autonomously take on open-ended research remains unclear.
To test it, they use Station, an open-world environment in which multiple agents simulate a scientific ecosystem. Open-ended tasks bring their own difficulty: there are often no intermediate metrics to tell an agent it is on the right track. The authors therefore add two mechanisms to Station, a Supervisor mechanism and periodic Meta Reflection, which are meant to encourage persistent exploration even when intermediate metrics are lacking.
The evaluation is built from three recent oral papers presented at ICLR. Each agent is given the main research question studied in the paper, while the paper's results are withheld and web access is disabled. The authors then measure how many of the original findings, partitioned into individual criteria, the agents rediscover.
On that measure Station rediscovers 62.7% of the criteria on average. The baselines score far lower: 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that adding the two mechanisms together improves research coverage and continuity.
The authors also run Station on two open-ended tasks that have no oracle paper. There, some of the discoveries made by the agents closely match discoveries reported by researchers after the knowledge cutoff date. From all this the authors conclude that a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery.
Key facts
- Station is an open-world environment where multiple agents simulate a scientific ecosystem; the authors add a Supervisor mechanism and periodic Meta Reflection to it.
- Tasks come from three recent ICLR oral papers: agents get the main research question, the results are withheld and web access is disabled.
- Station rediscovers 62.7% of the criteria on average, versus 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2.
- Ablation and behavioral analyses indicate the two mechanisms together improve research coverage and continuity.
- On two open-ended tasks without oracle papers, some of the agents' discoveries closely match discoveries researchers reported after the knowledge cutoff date.
Why it matters
Most evidence that AI can speed up science comes from settings with a clear score to push up. This paper targets the harder case, where an agent has a research question and no intermediate metric to steer by. Rediscovering withheld findings from published papers gives a concrete yardstick for that case. The headline gap is large in absolute terms: 62.7% of criteria for Station against 15.4% for Codex Multiagent-v2 and 14.4-20.6% for AI Scientist-v2. The authors read this as a sign that the environment, not only the model, shapes how far agents get in open-ended research.
Who it affects
Researchers building agentic research systems and automated-science tools are the direct audience, since the paper compares against two named baselines and proposes two mechanisms they could weigh against their own designs. Teams that benchmark research agents may also care about the evaluation recipe: take a published paper, give the agent only the research question, and score recovery of the original findings split into criteria.
How to use it
The abstract describes a recipe rather than a product. To evaluate an open-ended research agent the way the authors do, give it the main research question of a paper, withhold the results, disable web access, and count how many of the original findings, split into individual criteria, it rediscovers. To build a system like Station, the authors add a Supervisor mechanism and periodic Meta Reflection so that agents keep exploring when no intermediate metric is available.
How solid is it
The claims come from the authors' own abstract on Hugging Face Papers, so the numbers are self-reported. The comparison against two baselines is on tasks from just three ICLR oral papers. Ablation and behavioral analyses are reported to support the value of the two mechanisms together, but the abstract gives no figures for them. The two tasks without oracle papers are described only by the authors' statement that some discoveries closely match later researcher-reported ones.
Risks and caveats
The abstract names no authors or institutions, does not identify the three ICLR papers or the year, and does not say which language models power the agents. No per-task breakdown of the 62.7% average is given, and it does not say whether the average is over tasks or over pooled criteria. The matches on the two oracle-free tasks are only that some discoveries closely match later reports; there is no statement that the discoveries were verified or peer reviewed. No compute cost, run time or number of agents is given, so the comparison with baselines cannot be judged on cost.
“a suitable environment can enable agents to autonomously make meaningful progress in open-ended scientific discovery”
— Station paper abstract, Hugging Face Papers