OneSearch-VL unifies image and video deep research in one 8B agent

OneSearch-VL unifies image and video deep research in one 8B agent

A paper on Hugging Face introduces OneSearch-VL, a unified agent for deep research over single images, multiple images and video. The starting observation is that these three settings need different visual operations but share one workflow: visual grounding, external retrieval and fact composition. The hard part, the authors say, is preserving the dependencies that link localized visual anchors, entity relations, source-supported facts and the operations that produce the answer.

The agent is built around the Visually Grounded Evidence Graph (VGEG), which encodes those dependencies. The VGEG serves as a shared task-level reference for three jobs: data construction, process supervision and operation-level evaluation.

A VGEG-based data engine constructs and verifies multi-image and video questions and filters expert trajectories. From this data the authors assemble two training sets: OneSearch-VL-SFT-110K for supervised fine-tuning and OneSearch-VL-RL-10K for reinforcement learning. They also derive a reward from the VGEG annotations, the Evidence-aware Visual-Grounded Rubric reward (EVGR), which supervises evidence traceability and visual grounding during RL.

For fine-grained evaluation they build two new benchmarks, OneSearch-MI-Bench and OneSearch-Video-Bench, which organize questions by the research operations encoded in their VGEGs.

In experiments, OneSearch-VL-8B improves over Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively. The authors also report substantial gains across 7 image benchmarks and VideoDR. A project repository is linked at https://github.com/appletea233/OneSearch-VL.

Key facts

  • OneSearch-VL is one agent for single-image, multi-image and video deep research, centered on a Visually Grounded Evidence Graph (VGEG).
  • The VGEG links visual anchors, entity relations, source-supported facts and answer-producing operations, and is reused for data construction, process supervision and evaluation.
  • Training data comes in two sets, OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K, with an RL reward called EVGR that supervises evidence traceability and visual grounding.
  • Two new benchmarks, OneSearch-MI-Bench and OneSearch-Video-Bench, organize questions by research operation.
  • OneSearch-VL-8B beats Qwen3-VL-8B with tool access by 20.2 and 17.6 percentage points on the two new benchmarks, respectively, and reports substantial gains on 7 image benchmarks and VideoDR.

Why it matters

Deep research agents that answer questions by searching the web are usually discussed for text. This paper targets the visual side: questions about one image, several images or a video, where the agent must locate something in the visual input, retrieve outside sources and combine the facts. The authors' central claim is that these three settings share one workflow, so a single agent and a single evidence representation can cover them. Their VGEG is used at every stage, from building questions to rewarding the model during RL to scoring it, which keeps the data, the training signal and the evaluation tied to the same structure.

Who it affects

The work is aimed at researchers and engineers building multimodal search and research agents, especially those who need answers grounded in images or video and traceable to sources. It also gives evaluators two new benchmarks that organize questions by research operation rather than by a single score.

How to use it

The paper links a project repository at https://github.com/appletea233/OneSearch-VL. The training sets are named OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K, and the benchmarks are OneSearch-MI-Bench and OneSearch-Video-Bench. The source states no release date, license or weights availability; only a project repository link is given.

How solid is it

This is a paper abstract with the authors' own experimental claims; the headline numbers are 20.2 and 17.6 percentage points over Qwen3-VL-8B with tool access, reported on benchmarks the authors built themselves. The abstract names no authors or institutions. No absolute scores or baseline accuracies on any benchmark are given, only percentage-point improvements. The sizes of the two new benchmarks are not stated, and the abstract does not state the improvement magnitude on the 7 image benchmarks or on VideoDR, only 'substantial gains'.

Risks and caveats

Two of the headline gains come from benchmarks introduced in the same paper, so they have not been checked by outside evaluators. The only comparison named is Qwen3-VL-8B with tool access; no comparison with proprietary or other open models is mentioned. The abstract does not say which 7 image benchmarks are used. Without absolute scores, it is hard to judge how strong the starting point was or how large the final accuracy is.

“Single-image, multi-image, and video deep research require different visual operations but share a workflow of visual grounding, external retrieval, and fact composition.”

— OneSearch-VL paper abstract