OneStreamer streaming video LLM leads eight benchmarks it was compared on

OneStreamer streaming video LLM leads eight benchmarks it was compared on

The paper introduces OneStreamer, a streaming video LLM built around a problem the authors state plainly: a model watching a live video must keep evidence before it knows which future task that evidence will matter for, and must answer once enough evidence has arrived. The hard part, they say, is forming reusable factual memory without hurting real-time perception.

OneStreamer handles this by learning two things together: query-independent evidence recording (memory) and task response. Both run through one shared proactive generation process. The memory component is called Proactive Hierarchical Caption Memory (PHCM). It produces time-grounded captions of local detail, plus summaries of events that have completed. During training, streaming caption targets supervise how the model interprets the video prefixes it has already observed. At inference, the captions the model generated itself sit alongside a recent visual window, so they supply reusable factual context without the model having to revisit historical visual features.

A second component, Proactive State Transition Learning (PSTL), deals with a training imbalance. A streaming model spends long stretches in repeated waiting states, and these dominate the supervision. PSTL reduces that dominance by keeping supervision at all output anchors and selecting representative state-change and state-persistence tokens.

The authors also built a streaming data synthesis pipeline that aligns output content and timing with the evidence actually available. They combined the streaming captions and QA it produced with cleaned open-source data to make OneStreamer-1M, a streaming video interaction dataset with over one million records spanning diverse tasks.

On results, the authors report that their 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining the generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. The authors conclude that these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.

Key facts

  • OneStreamer jointly learns query-independent evidence recording and task response through one shared proactive generation process.
  • Proactive Hierarchical Caption Memory (PHCM) writes time-grounded local-detail captions and summaries of completed events; at inference these complement a recent visual window instead of revisiting historical visual features.
  • Proactive State Transition Learning (PSTL) beats dense state supervision while supervising only 27.5% of annotated state tokens.
  • The 4B model achieves the best results among the compared methods on all eight evaluated streaming video understanding benchmarks.
  • The new OneStreamer-1M dataset has over one million records spanning diverse tasks, built from synthesized streaming captions and QA plus cleaned open-source data.

Why it matters

Streaming video is hard for language models because the question often arrives after the footage. The model has to decide what to remember before it knows what will be asked. OneStreamer's answer is to have the model write its own time-grounded captions and event summaries as it watches, and to train recording and responding as one process rather than two. The authors present this as a way to get reusable factual memory without hurting real-time perception.

Who it affects

The work is aimed at researchers building streaming video understanding systems, where a model must follow a continuing video and answer at the right moment. The OneStreamer-1M dataset, with over one million records across diverse tasks, is also a training resource for that line of work.

How to use it

The abstract describes a method and a dataset rather than a product. It gives no information on whether code, weights or the OneStreamer-1M dataset are released. For practitioners, the transferable ideas are the ones described: keep generated captions as memory next to a recent visual window, and use PSTL-style selection of state-change and state-persistence tokens instead of dense state supervision.

How solid is it

The claims come from the authors' own abstract. The headline result is limited to the methods the authors chose to compare, and the compared set is not listed. No benchmark names, baseline names, scores or margins are given, and the size of the historical QA improvement is not stated. Ablations back two specific points: retaining generated captions improves historical QA without degrading real-time perception, and PSTL outperforms dense state supervision with 27.5% of annotated state tokens.

Risks and caveats

Without scores or baseline names, the strength of the lead cannot be judged from this text. The base model of the 4B model is not named, and no latency, throughput, memory or compute figures are given, so the cost of generating captions during streaming is unknown. The authors say their results support proactive generation as a shared learning interface; that is framed as support, not proof.

“Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available.”

— OneStreamer paper abstract