動画AIの途中層にある時間情報を出力まで保つ
Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
この論文をやさしく読む
ひとことで言うと
動画AIが途中で捉えた出来事の順序情報を、後の層へ渡し直して回答に反映させる方法です。
何に役立つ?
動画中の前後関係を問う課題で、追加訓練なしに推論を改善する用途が示されています。
この研究の面白いところ
最初から時間情報を得られていないのではなく、途中で得た差が後の層で弱まることに着目しています。
どこまで分かった?
改善は3モデル・4ベンチマークでの評価です。要旨に具体的な改善率や追加計算量はなく、すべての動画課題に対する保証は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動画大規模言語モデル(VideoLLM)は、フレームを順番に受け取り、視覚内容が時間軸に沿ってどう変化するかを解釈するが、時間的推論は構造の違いを超えて持続的な弱点となっている。動画のフレーム順序を反転すると、時間関係に関する答えも反転するはずなのに、最終的な予測は変わらないことが多い。この失敗の発生源を調べるため、時間順序の反転によって生じる各層の表現の差を、時間的差異ベクトルτ_lとして定義する。その大きさを層ごとに追うと、差異が中間層で最大となり、出力へ向かって徐々に減少する、一貫した時間的差異のプロファイルが現れる。このピークが時間的推論に特有で、予測にとって機能的に重要であることを確認し、VideoLLMが中間層で時間情報を獲得しても、出力まで維持できないことを示す。この徐々に薄れる性質を踏まえ、入力ごとにプロファイルのピークでτ_lを取り出し、測定した減衰に沿って後続の層へ再注入するTemporal Activation Injection(TAI)を提案する。TAIは訓練を必要とせず、三つのVideoLLMと四つのベンチマークにわたって時間的推論を一貫して改善し、時間関係を扱わない課題への影響はごく小さい。コードは https://github.com/Youngwoo-git/Before-It-Fades で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves the final prediction unchanged. We investigate where this failure originates by defining the temporal divergence vector $\tau_l$, the layer-wise representational difference induced by reversing temporal order. Tracking its magnitude across layers reveals a consistent temporal divergence profile where the divergence peaks at intermediate layers and progressively diminishes toward the output. We confirm this peak is specific to temporal reasoning and functionally critical for predictions, establishing that VideoLLMs acquire temporal information at intermediate layers but fail to maintain it to the output. This progressive fading motivates our method, Temporal Activation Injection (TAI), which extracts $\tau_l$ at the peak of the profile for each input and reinjects it into subsequent layers following the measured decay. TAI requires no training and consistently improves temporal reasoning across three VideoLLMs and four benchmarks with negligible impact on non-temporal tasks. Code is available at https://github.com/Youngwoo-git/Before-It-Fades.
著者のコメント
Accepted to NeurIPS 2026
arXiv ID: 2610.01595 / 要約の誤りについて