Temporal Activation Injection (TAI) restores fading time cues in VideoLLMs without training

Temporal Activation Injection (TAI) restores fading time cues in VideoLLMs without training

Video Large Language Models (VideoLLMs) read frames in sequence and are meant to interpret how visual content changes over time. Yet the paper says temporal reasoning remains a persistent weakness across architectures. One telling symptom: reversing the frame order of a video, which should invert temporal answers, often leaves the model's final prediction unchanged.\n\nTo find where the failure starts, the authors define a temporal divergence vector, written τ_l: the layer-wise representational difference induced by reversing temporal order. They track its magnitude layer by layer and see a consistent temporal divergence profile. The divergence peaks at intermediate layers and then progressively diminishes toward the output. The authors confirm that this peak is specific to temporal reasoning and functionally critical for predictions. Their conclusion is that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it all the way to the output.\n\nThis progressive fading motivates the method, Temporal Activation Injection (TAI). For each input, TAI extracts the divergence vector at the peak of the profile and reinjects it into subsequent layers, following the decay that was measured. TAI requires no training. According to the authors, it consistently improves temporal reasoning across three VideoLLMs and four benchmarks, with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.

Key facts

  • Reversing a video's frame order, which should invert temporal answers, often leaves a VideoLLM's final prediction unchanged.
  • The temporal divergence vector (the layer-wise representational difference caused by reversing temporal order) peaks at intermediate layers and fades toward the output.
  • The authors confirm the peak is specific to temporal reasoning and functionally critical for predictions, so the model acquires temporal information but fails to maintain it.
  • Temporal Activation Injection (TAI) extracts the vector at the peak for each input and reinjects it into later layers following the measured decay; it requires no training.
  • TAI is reported to improve temporal reasoning consistently across three VideoLLMs and four benchmarks, with negligible impact on non-temporal tasks.

Why it matters

Temporal reasoning is described as a persistent weakness across VideoLLM architectures, and the frame-reversal test shows how deep the problem goes: a model can answer the same way whether a video plays forward or backward. The paper's contribution is a diagnosis as much as a fix. It locates the failure inside the network: temporal information is built up at intermediate layers and then fades before it reaches the output. That framing turns a vague weakness into a measurable layer-by-layer pattern.

Who it affects

Researchers and engineers working with VideoLLMs, especially anyone evaluating or improving how these models handle order and change over time. Because the method is training-free, it also concerns people who use existing VideoLLMs and cannot retrain them. The paper evaluates three VideoLLMs and four benchmarks.

How to use it

TAI requires no training, so it is applied at inference time to an existing model. For each input it takes the temporal divergence vector at the peak layer and adds it back into subsequent layers following the measured decay. Code is available at https://github.com/Youngwoo-git/Before-It-Fades. The source states no inference cost or latency overhead, so that is something to measure before relying on it.

How solid is it

The claims come from the authors' own abstract. They report that the divergence peak is specific to temporal reasoning and functionally critical for predictions, and that TAI consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. The source does not name the three VideoLLMs or the four benchmarks, and gives no quantitative improvement figures such as accuracy gains. The layer index of the peak and the decay form are not specified either. Treat the size of the benefit as unknown until the full paper is checked.

Risks and caveats

No limitations or failure cases are described in the source. The size of the gains is not given, so the improvement may be small or large; only its consistency and the near-absence of harm to non-temporal tasks are claimed. The results cover three VideoLLMs and four benchmarks, so how far they extend to other models is not established here. No inference cost or latency overhead is stated.

“Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged.”

— From the paper's abstract