OmniCapBench tests audio-visual captioning across 786 videos

A paper introduces OmniCapBench, short for Omni-Video Caption Benchmark, an evaluation framework for audio-visual captioning by multimodal large language models (MLLMs). The authors start from the observation that MLLMs are moving toward continuous audio-visual reasoning, which creates a need for evaluations that expose where their capabilities run out. Captioning audio and video together is, in their view, a good diagnostic task, but existing benchmarks face a coupled trade-off: whole-caption scores give coverage without localization, local probes give localization without coverage, and unconstrained LLM judges introduce instability.
OmniCapBench reframes caption evaluation as a deep-structured diagnostic framework. Instead of grading free-form text, it shifts the prediction target to sets of atomic, verifiable evaluation units. These are organised into three tracks: entity references, visual shots, and audio events. Scoring combines deterministic constraint checks with localized LLM-based semantic comparisons, which the authors say makes scoring reliable.
The benchmark rests on 786 densely annotated videos. According to the authors, it effectively distinguishes several kinds of MLLM perception error: temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions.
When the authors evaluate frontier MLLMs, they find strong local perception but weak long-horizon audio-visual reasoning, particularly in identity drift and cross-modal misalignment. They present this as a fine-grained roadmap for omnimodal development.
Key facts
- OmniCapBench (Omni-Video Caption Benchmark) evaluates audio-visual captioning using 786 densely annotated videos.
- Instead of free-form text, models are scored on sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events.
- Scoring combines deterministic constraint checks with localized LLM-based semantic comparisons, aiming to avoid the instability of unconstrained LLM judges.
- The benchmark is said to separate temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions.
- Frontier MLLMs show strong local perception but weak long-horizon audio-visual reasoning, especially on identity drift and cross-modal misalignment.
Why it matters
Captioning video with sound is a demanding test of whether a model can follow who is who, what is seen and what is heard over time. The authors argue that current benchmarks force a choice: whole-caption scores cover everything but cannot say where a model went wrong, local probes pinpoint errors but miss the overall picture, and unconstrained LLM judges are unstable. OmniCapBench tries to remove that trade-off by breaking a caption into atomic, verifiable units and scoring each one, so a failure can be traced to a specific track and error type.
Who it affects
The framework is aimed at people building and evaluating omnimodal and multimodal language models, since it is meant to expose capability limits and guide development. The paper reports on frontier MLLMs in general rather than naming particular models.
How to use it
The benchmark is organised around three tracks: entity references, visual shots, and audio events. A model's caption is judged against sets of atomic units, with deterministic constraint checks handling what can be verified mechanically and localized LLM-based comparisons handling semantic matches. The authors present the resulting error breakdown as a roadmap for omnimodal development. The abstract states no release date, dataset license, or code and data availability.
How solid is it
This is a preprint abstract, and every finding in it is the authors' own claim. The abstract gives the size of the benchmark (786 densely annotated videos) and the structure of the three tracks, but names no authors or institutions, no specific models evaluated, and no scores, rankings or numeric results for any model. The headline finding about frontier MLLMs therefore cannot be checked from the abstract alone.
Risks and caveats
The evaluation still relies on an LLM for the localized semantic comparisons, and the abstract does not say which LLM is used. Claims that the benchmark effectively distinguishes error types and offers a fine-grained roadmap come from the authors, without numeric results in the abstract to support them. The abstract also does not say how many evaluation units or annotators there are, or the total video duration.
“a fine-grained roadmap for omnimodal development”
— OmniCapBench abstract