EgoTools introduces 100-hour video dataset and benchmark for tool-use reasoning

EgoTools introduces 100-hour video dataset and benchmark for tool-use reasoning

A paper titled "EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos" targets a gap in video AI. Real-world embodied tasks, from everyday activities to professional procedures, require agents to act under physical constraints while tracking changing object and task states. Many of these activities are tool-mediated, and understanding them means reasoning about affordances, hand-tool-object geometry, procedural progress and the causal effects of a tool on its target object. The authors say that despite strong results on perception-oriented video tasks such as captioning and general video QA, current multimodal video models remain limited in this kind of tool-centric embodied reasoning. They add that progress has been held back by a lack of real-world egocentric data and diagnostic benchmarks.

To fill that gap they introduce EgoTools, which they call the first comprehensive suite for egocentric tool-use understanding. It has two parts. EgoTools-Data is a large-scale corpus of 100 hours of tool-centric egocentric recordings, with synchronized audio, dense captions, reasoning-heavy narrations and supplementary 3D information. EgoTools-Bench is a diagnostic benchmark of 1,000 QA pairs across four tracks, covering tool-use understanding from perception and geometry to procedure and causal reasoning.

The experiments, according to the authors, show that current models still struggle to ground tool use in visual evidence. Gemini-3.1-Pro reaches 66.9% overall accuracy on the benchmark but only 51.7% on the Perception & Grounding track, a gap of 15.2 percentage points.

The authors also test EgoTools-Data as a training resource. On the full 1,000-question benchmark, full supervised fine-tuning lifts Qwen3-VL-8B-Instruct from 50.0% to 60.9% accuracy, a gain of 10.9 percentage points. They state that this is measured under strict source-video separation. Their conclusion is that EgoTools works as a unified resource for both training and diagnostic evaluation of real-world egocentric tool-use understanding.

Key facts

  • EgoTools is presented by its authors as the first comprehensive suite for egocentric tool-use understanding, with a training corpus and a diagnostic benchmark.
  • EgoTools-Data holds 100 hours of tool-centric egocentric recordings with synchronized audio, dense captions, reasoning-heavy narrations and supplementary 3D information.
  • EgoTools-Bench has 1,000 QA pairs across four tracks, spanning perception and geometry through procedure and causal reasoning.
  • Gemini-3.1-Pro scores 66.9% overall on the benchmark but only 51.7% on Perception & Grounding.
  • Full supervised fine-tuning on EgoTools-Data raises Qwen3-VL-8B-Instruct from 50.0% to 60.9% on the full benchmark, under strict source-video separation.

Why it matters

The authors argue that models do well on captioning and general video QA yet remain limited when a task hinges on how a tool is used: its affordances, the geometry between hand, tool and object, how far a procedure has progressed, and what the tool does to its target. They say the missing pieces were real-world egocentric data and diagnostic benchmarks. EgoTools is their attempt to supply both. The Gemini-3.1-Pro result, 66.9% overall but 51.7% on Perception & Grounding, is the clearest sign in the paper that grounding tool use in visual evidence is where models fall short.

Who it affects

The material is aimed at people building and evaluating multimodal video models, and at work on embodied agents that must act under physical constraints and track object and task states. The abstract frames tool-mediated activities as spanning everyday tasks and professional procedures.

How to use it

The suite is meant for two purposes: EgoTools-Data as training material and EgoTools-Bench as a diagnostic test. The authors show the training use with Qwen3-VL-8B-Instruct, fine-tuned in a full supervised setup. No release date, license, or download link for the data or benchmark is given in the text.

How solid is it

The claims come from the authors' own abstract. The fine-tuning gain is reported on the full 1,000-question benchmark under strict source-video separation, which the authors cite as a safeguard in the measurement. The headline claim of being the first comprehensive suite is the authors' own framing. Scores of models other than Gemini-3.1-Pro and Qwen3-VL-8B-Instruct are not given, and the full supervised fine-tuning setup (epochs, hardware, data subset) is not described.

Risks and caveats

Only two models are reported with numbers, so the picture of how models in general handle tool use is partial. Names of the four benchmark tracks other than Perception & Grounding are not given, and neither are Gemini-3.1-Pro scores on those tracks. The 10.9-point gain for Qwen3-VL-8B-Instruct is measured on the same benchmark the corpus was built alongside, so it shows the data is useful for this benchmark, which is what the authors set out to validate.