LoHi trades frame resolution for density, lifting long-video VLM accuracy

A paper listed on Hugging Face Papers argues that efficient long-video understanding with vision-language models (VLMs) is usually framed too narrowly. The common approach is to pick informative frames or visual tokens at a fixed native resolution. The authors show that per-frame resolution can instead be traded for denser temporal coverage, and that front-end decoding latency depends on the size of the candidate pool rather than the final token budget.
The argument rests on an empirical study across multiple VLMs and long-video benchmarks, which yields three findings. First, dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets. Second, resolution-sensitive tasks benefit from selected high-resolution frames. Third, front-end decoding dominates wall time for hour-long videos.
Building on these findings, the paper introduces LoHi, a training-free, single-pass framework. It combines a dense low-resolution video stream with sparse high-resolution image frames, and feeds both through the VLM's native video and image pathways. Two variants decide which frames get full resolution. LoHi-Anchor selects them using codec I-frame metadata. LoHi-SemDiv uses query relevance and visual diversity over CLIP features.
Across three long-video benchmarks, LoHi improves average accuracy by 10.6 percentage points over the native-resolution baseline at a matched token budget. It beats the strongest prior efficiency method by 5.2 percentage points. It also reduces front-end decoding latency by up to 7x on hour-long videos. A project page is linked from the listing.
Key facts
- LoHi is a training-free, single-pass framework that combines a dense low-resolution video stream with sparse high-resolution image frames, using the VLM's native video and image pathways.
- Across three long-video benchmarks, average accuracy is 10.6 percentage points above the native-resolution baseline at a matched token budget, and 5.2 points above the strongest prior efficiency method.
- Front-end decoding latency falls by up to 7x on hour-long videos.
- Two frame selectors are offered: LoHi-Anchor uses codec I-frame metadata, and LoHi-SemDiv uses query relevance and visual diversity over CLIP features.
- The motivating study found that dense low-resolution sampling beats sparse native-resolution sampling at equal token budgets, and that decoding dominates wall time on hour-long videos.
Why it matters
Long videos are expensive for VLMs, and most efficiency work asks which frames or tokens to keep at native resolution. This paper changes the question: spend the same token budget on more frames at lower resolution, and keep full resolution only for a few frames that need it. It also points at a cost that is easy to overlook. For hour-long videos, front-end decoding dominates wall time, and that latency depends on the candidate pool size, not on the final token budget.
Who it affects
Researchers and engineers building long-video question answering or video analysis on top of existing VLMs. Because LoHi is training-free and works through the model's native video and image pathways, it is aimed at people who want to improve a model they already use rather than train a new one. Teams who process hour-long footage and care about decoding time are the most direct audience.
How to use it
The abstract describes the approach but gives no setup instructions. In outline: feed the VLM a dense low-resolution video stream, add a small set of high-resolution frames as images, and choose those frames with one of two selectors. LoHi-Anchor relies on codec I-frame metadata, while LoHi-SemDiv scores frames by query relevance and visual diversity using CLIP features. The paper links a project page at https://sixundong.com/projects/lohi.
How solid is it
The listing carries the abstract only, so the numbers cannot be checked against tables here. The reported gains are averages across three long-video benchmarks, in percentage points: 10.6 over the native-resolution baseline at a matched token budget and 5.2 over the strongest prior efficiency method. They are not per-benchmark results and not relative improvements. The 7x latency reduction is an upper bound ('up to') for hour-long videos. The abstract does not name the benchmarks or the VLMs tested, nor the token budget, resolutions or frame counts.
Risks and caveats
The headline figures are averages, and the 7x speedup is a best case, so typical results on other models or video lengths may differ. The abstract does not name the benchmarks, the models or the budgets, which limits how far the numbers can be generalised before reading the full paper. It also says that resolution-sensitive tasks need selected high-resolution frames, so the choice of selector matters for those tasks.