Incremental-OEDR updates deep research reports instead of rewriting them

Incremental-OEDR updates deep research reports instead of rewriting them

A new paper takes aim at how Open-Ended Deep Research (OEDR) systems work today. According to the authors, these systems mostly generate a report from scratch each time, which is inefficient when a report has to be kept current as new information emerges.

The paper introduces Incremental Open-Ended Deep Research, or Incremental-OEDR. In this setting a report is treated as an evolving research state and is updated step by step: valid knowledge is preserved, outdated or incomplete content is revised, and newly available information is incorporated.

To support the setting, the authors propose a method called Structured Harness. It represents a report as a structured collection of outlines, sections and supporting evidence. It provides three things: structured retrieval, a persistent structured evidence pool, and structured generation, which together allow selective report updating and reuse of evidence already gathered.

The authors also build a temporal evaluation framework spanning ten years. It has two tasks: a Single-Step Task, which evaluates individual transitions from one report state to the next, and a Long-Chain Task, which evaluates long-term chains of updates.

Experiments were run on DeepResearch Bench and DeepConsult, under both an Open-source Configuration (OC) and a Proprietary Configuration (PC). The authors report that Incremental-OEDR keeps competitive report quality while substantially improving report continuity and reducing research costs. On DeepResearch Bench, they state it reaches up to 0.51 higher content-level ROUGE-L F1 and 0.63 higher outline-level EM F1 than OEDR, with 33% lower token consumption and 61% fewer search calls. A project page at https://ioedr-project.github.io/ carries more details.

Key facts

  • Incremental-OEDR treats a deep research report as an evolving state and updates it, rather than regenerating it from scratch each time new information appears.
  • The Structured Harness method stores reports as outlines, sections and supporting evidence, with structured retrieval, a persistent evidence pool and structured generation for selective updates.
  • On DeepResearch Bench the authors report up to 0.51 higher content-level ROUGE-L F1 and 0.63 higher outline-level EM F1 than OEDR.
  • The same comparison shows 33% lower token consumption and 61% fewer search calls than OEDR.
  • The evaluation framework spans ten years and has a Single-Step Task and a Long-Chain Task; experiments also cover DeepConsult under open-source and proprietary configurations.

Why it matters

Most deep research tools produce a finished report once and start over when the topic moves. The paper frames report maintenance as its own problem: a report that must stay current as new information emerges. Treating the report as persistent state, with a reusable evidence pool, is the central idea, and the claimed payoff is lower cost (fewer tokens and search calls) without giving up report quality.

Who it affects

The work is aimed at people building or evaluating Open-Ended Deep Research systems, especially for scenarios where research reports need to be continuously maintained. It also offers a benchmark-style evaluation setup, with single-step and long-chain update tasks, that other researchers could use to test incremental updating.

How to use it

The paper points to a project page at https://ioedr-project.github.io/ for more details. The method itself is described as a harness built around structured retrieval, a persistent structured evidence pool and structured generation, so a team would organise a report as outlines, sections and linked evidence and update only the parts that changed.

How solid is it

This is a preprint and the figures are the authors' own results on DeepResearch Bench, with experiments also run on DeepConsult under both an Open-source Configuration and a Proprietary Configuration. The headline gains are stated as 'up to' values. The abstract gives no absolute baseline values for OEDR, so the starting points of the gains are unknown, and it does not say which configuration the 'up to' figures come from. No numeric results are given for DeepConsult, and 'competitive report quality' is not quantified.

Risks and caveats

The reported ROUGE-L F1 and EM F1 gains are score differences, not percentages of a baseline, and the best-case 'up to' wording means typical gains may be smaller. The abstract does not name the underlying models used in either configuration, and cost savings are expressed only as token use and search calls, not money. It is not stated what the ten-year span of the temporal evaluation consists of. The claims have not been independently checked here.

“treats a report as an evolving research state and incrementally updates it by preserving valid knowledge, revising outdated or incomplete content, and incorporating newly available information”

— Incremental-OEDR paper abstract, describing the setting