AutoSciBench generates and revises scientific-agent benchmarks automatically

AutoSciBench generates and revises scientific-agent benchmarks automatically

The paper starts from a problem with agent evaluation. As agents evolve quickly, existing benchmarks can become saturated, which limits their ability to tell capabilities apart and to expose the failure modes that remain. In scientific domains the problem is worse, because building and updating a benchmark takes substantial time, labor and domain expertise. That makes it hard to keep evaluation in step with what agents can do. The authors ask whether scientific-agent benchmarks can instead be generated automatically and adapted iteratively as agent capabilities change.

Their answer is AutoSciBench. The framework represents every task in two layers. A high-level concept specifies the scientific domain, the data modality and the required reasoning approach. A low-level recipe specifies how the question, the environment and the ground-truth answer are constructed and verified.

The refinement loop works like this. Agents attempt each task, which produces solver trajectories and matching judge feedback. AutoSciBench uses these to revise the recipe or the concept. The revisions close shortcuts that the solvers were observed to take, and they push tasks toward raw-data re-examination, interpretation of intermediate results and evidence integration. A further step reuses what was learned: experience distilled from completed refinement trajectories guides the generation of new concepts, so lessons from earlier task refinement inform later benchmark construction.

For the evaluation, the authors started from existing benchmarks and tested AutoSciBench across computational biology, materials science and clinical imaging. The generated benchmarks reduced average solver accuracy by 22.4 percentage points in computational biology and 25.5 percentage points in materials science, relative to the human-curated benchmarks. These are absolute differences in percentage points, not relative percentages. Generated tasks also received higher average quality ratings across all three domains. The authors conclude that these results suggest scientific-agent evaluation can adapt as agent capabilities advance.

Key facts

  • AutoSciBench is a framework that automatically generates scientific-agent benchmarks and iteratively revises them as agent capabilities evolve.
  • Each task has a high-level concept (scientific domain, data modality, required reasoning approach) and a low-level recipe (how the question, environment and ground-truth answer are built and verified).
  • Solver trajectories and judge feedback drive revisions that close observed shortcuts and shift tasks toward raw-data re-examination, interpretation of intermediate results and evidence integration.
  • Generated benchmarks cut average solver accuracy by 22.4 percentage points in computational biology and 25.5 in materials science, relative to human-curated benchmarks.
  • Generated tasks received higher average quality ratings across all three tested domains: computational biology, materials science and clinical imaging.

Why it matters

Benchmarks that agents have saturated stop telling researchers which systems are better or where they still fail. Refreshing them by hand in science is slow, because it takes time, labor and domain expertise. AutoSciBench tests a different route: let the benchmark be rebuilt automatically as agents improve. The reported accuracy drops of 22.4 and 25.5 percentage points against human-curated benchmarks show the generated versions were harder for the solvers in two of three domains, and the authors say this suggests evaluation can keep pace with agent progress.

Who it affects

Researchers who build and maintain benchmarks for scientific agents are the most direct audience, since the framework targets the manual effort of constructing and updating them. Teams that evaluate agents in computational biology, materials science and clinical imaging, the three domains tested, are next in line.

How to use it

The source text describes a method, not a product. It covers the concept and recipe task representation, the solve, judge and revise loop, and the reuse of distilled experience for new concepts. No code, dataset release or compute cost is mentioned. A practitioner can take away the design pattern: describe tasks at two levels, run agents on them, and feed trajectories and judge feedback back into the task definition.

How solid is it

The evidence is the paper's own abstract. It reports two accuracy reductions, 22.4 points in computational biology and 25.5 in materials science, plus higher average quality ratings in all three domains. The authors phrase the conclusion as a suggestion, not a proof. The accuracy reduction is reported only for computational biology and materials science; no accuracy figure is given for clinical imaging. Which agents or models served as solvers or judges is not stated. Neither is how the quality ratings were produced, or their scale.

Risks and caveats

Several details needed to judge the result are not stated. Absolute accuracy levels of solvers on either benchmark type are not given, the names of the human-curated benchmarks are not stated, and the size of the generated benchmarks (number of tasks) is not stated. A lower solver accuracy shows the generated tasks were harder for the tested solvers, and the quality ratings are the paper's separate check on that. Because the evaluation started from existing benchmarks in three domains, the results say little about other fields.