動画を見ながら記録を作り、適切な時点で応答するOneStreamer
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
この論文をやさしく読む
ひとことで言うと
流れ続ける動画について、後で役立つ事実を文章で記憶し、答えられる証拠がそろったら応答するモデルです。
何に役立つ?
過去の出来事を尋ねながら現在の映像も扱う動画対話に役立つ可能性があります。8ベンチマークでは、比較した手法の中で最良の結果が報告されています。
この研究の面白いところ
質問が来る前から記録を作ることと、応答のタイミングを学ぶことを同じ生成過程で扱います。直近の映像と過去の文章記憶を組み合わせる設計です。
どこまで分かった?
27.5%は教師信号を与えた注釈付き状態トークンの割合であり、全学習計算量の割合ではありません。要旨には各ベンチマークの具体的スコアや実運用時の応答遅延は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ストリーミング動画を扱うLLMは、将来のタスクとの関連性がまだ分からない段階で証拠を保持し、十分な証拠が得られた時点で応答しなければならない。課題は、リアルタイムの知覚を損なわず、再利用できる事実の記憶を作ることである。本研究では、共通の能動的な生成過程を通じて、質問に依存しない証拠の記録とタスクへの応答を同時に学習するOneStreamerを導入する。 能動的階層キャプション記憶(PHCM)は、時刻に対応付けられた局所的な詳細のキャプションと、完了した出来事の要約を生成する。学習時には、ストリーミングキャプションの教師信号により、その時点までに観測した動画の解釈を学習する。推論時には、モデルが生成した記録が直近の映像区間を補い、過去の視覚特徴を再参照せずに、再利用可能な事実の文脈を提供する。 能動的状態遷移学習(PSTL)は、すべての出力アンカーにおける教師信号を保持し、状態の変化と維持を表す代表的なトークンを選ぶことで、繰り返される待機状態の過度な支配を抑える。さらに、出力内容と出力時点を利用可能な証拠に整合させる、ストリーミングデータの合成パイプラインを開発する。得られたストリーミングキャプションと質問応答を、整備したオープンソースデータと組み合わせることで、多様なタスクにわたる100万件超の記録を持つ、広範囲のストリーミング動画対話データセットOneStreamer-1Mを構築した。 40億パラメータの本モデルは、評価したストリーミング動画理解の8ベンチマークすべてで、比較手法中の最良の結果を達成した。構成要素の比較実験では、生成キャプションを保持すると、リアルタイム知覚を損なわずに過去についての質問応答が改善した。PSTLは、注釈付き状態トークンの27.5%だけを教師信号の対象としながら、密な状態教師信号を用いる方法も上回った。これらの結果は、能動的な生成を、ストリーミング動画対話における知覚、記憶形成、適時の応答をつなぐ共通の学習インターフェースとする考えを支持する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
著者のコメント
29 pages, 12 figures, 20 tables. Project page: https://mcg-nju.github.io/OneStreamer
arXiv ID: 2610.01762 / 要約の誤りについて