arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

自己回帰型動画生成のための記憶機構を整理

The Past Frames the Future: Memory for Autoregressive Video Generation

Harold Haodong Chen, Rongjin Guo, Disen Lan, Wen-Jie Shu, Hongfei Zhang, Hanzhe Hu, Shengtao Yao, Zixin Zhang, Guibin Zhang, Zhefan Rao, Jinxiu Liu, Yexin Liu, Rui Peng, Yuhao Liu, Bin Ren, Shuai Yang, Yukang Chen, Salman Khan, Ying-Cong Chen, Ser-Nam Lim, Rynson W.H. Lau, Nicu Sebe, Yu Cheng, Ming-Hsuan Yang, Qifeng Chen

この論文をやさしく読む

ひとことで言うと

長い動画を順に作るモデルで、過去の情報を保持・更新・評価する方法を整理したレビューである。

何に役立つ?

考えられる用途は、長時間の動画生成で人物や状態の一貫性を保つ仕組みの設計である。要旨は新たな性能実験ではなく文献の整理を説明する。

この研究の面白いところ

記憶を、元の場面が文脈窓から消えても将来の生成に効く情報と定義し、五つの観点で研究を整理した。

どこまで分かった?

未解決課題として資源制約、状態更新、評価の標準化が挙げられる。特定方式の有効性を実証したとの記述はない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

生成モデルの進歩により動画の忠実度が向上し、長い時間にわたる生成、対話的な世界モデル、変化し続ける視覚環境が可能になってきた。自己回帰型の動画生成は、因果的な展開によって映像系列を延長する。しかし、系列が長くなると、実用的なモデルには文脈窓、保存容量、計算量の厳しい上限がある。その結果、登場するものの同一性、動的な状態、介入で生じた因果的変化などの重要な過去情報が、まだ必要なうちに現在の文脈から外れてしまう。これを克服し、時間を通じた持続性を保つことは、記憶の根本的な問題である。 本論文は、自己回帰型動画生成における記憶機構を体系的かつ包括的にレビューする。記憶を、外側の自己回帰ステップをまたいで保持され、元の証拠が現在の局所的な文脈から消えた後も将来の生成に影響できる過去情報として、操作的に定義する。この共通の枠組みに基づき、既存研究を五つの相補的な観点から整理する。第一は履歴を担う表現形式、第二は保持すべき意味情報と物理情報という機能、第三は記憶の書き込み、読み出し、更新、管理、統合という操作、第四は閉ループの展開を通じた記憶動作の学習、第五は本当の記憶能力を診断する評価である。 最後に、組み合わせ可能で資源を意識した記憶構造、信頼できる状態更新、自己生成した展開を使う学習、標準化された評価などの未解決課題をまとめる。表現、仕組み、学習の方法を結び付けることで、信頼性の高い記憶条件付き動画生成システムを開発するための整理された基盤を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Advances in generative models have improved video fidelity, enabling long-horizon generation, interactive world modeling, and evolving visual environments. Autoregressive (AR) video generation extends visual sequences through causal rollouts. However, a fundamental bottleneck emerges: as the generated sequence expands, practical models must operate under strictly bounded context windows, storage, and computational limits. Consequently, critical historical information, e.g., entity identities, dynamic states, and intervention-induced causal changes, often leaves the active context long before its relevance diminishes. Overcoming this limitation and maintaining temporal persistence constitutes a fundamental memory problem. We present a systematic and comprehensive review of memory mechanisms in AR video generation. We formulate memory operationally as persistent historical information maintained across outer AR steps, capable of influencing future generation even after the originating evidence is no longer locally accessible. Building upon this unified framework, we organize the literature through five complementary perspectives: (I) Forms, the representational carriers of history; (II) Functions, the specific semantic and physical information requiring preservation; (III) Operations, the lifecycle of writing, reading, updating, managing, and integrating memory; (IV) Learning, the optimization of memory behaviors under closed-loop rollouts; and (V) Evaluation, the paradigms for diagnosing genuine memory capabilities. We conclude by synthesizing open challenges, including composable and resource-aware memory architectures, trustworthy state updating, self-rollout learning, and standardized evaluation. By bridging representations, mechanisms, and learning paradigms, this paper establishes a structured foundation for developing reliable, memory-conditioned video generation systems.

arXiv ID: 2609.28466 / 要約の誤りについて