arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

生成動画の世界で新しい物体への操作と状態の持続を両立

Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models

Xindi Yang, Baolu Li, Liam Lee, Zhenfei Yin, Songxin Zhang, Zhuoyang Song, Xu Jia, Jianfei Cai, Tien-Tsin Wong, Bingyi Jing, Mengyue Yang

この論文をやさしく読む

ひとことで言うと

動画で生成した世界の物体や操作結果を別の状態表に記録し、新しい物体にも操作でき、過去の変更も残るようにする方法です。

何に役立つ?

探索中に内容が増える仮想環境で、操作の結果が次の場面にも反映されるシミュレーションや対話環境への利用が考えられます。

この研究の面白いところ

映像だけに記憶を任せず、エージェントが更新する世界状態を中間に置き、その状態から動画を再生成しています。

どこまで分かった?

要旨は新規物体との相互作用と長時間の状態持続を実験で示したとしていますが、具体的な時間長、評価指標、失敗条件は記載していません。物理的な正確さ全般の保証とも述べられていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

生成型の動画世界モデルは、エージェントが移動し、単純な形で相互作用できる、開かれた環境を合成できるようになった。しかし、制限のない生成は、十分な相互作用を意味しない。生成世界が広がるにつれ、移動によって新たに作られた内容はエージェントが働きかけられる対象を増やすべきであり、エージェントが世界を変えたとき、その変化は一時的な視覚効果ではなく、環境の持続的な部分になるべきである。 本研究では、この二つの要件を、新たに生成または遭遇した実体が操作可能な世界へ取り込まれる「開世界の相互作用性」と、相互作用の結果が世界状態に確定され、その後の観測や相互作用へ影響し続ける「持続する状態」として特徴付ける。コーディングエージェントが管理する明示的かつ拡張可能な世界状態を通じ、生成と相互作用を循環させる対話的な動画世界モデルOneiraを提示する。 現在の観測と行動または高水準の目標が与えられると、エージェントは世界状態を読み、関係する実体を特定し、相互作用を計画して、その結果を世界状態表へ書き戻す。探索により新しい物体が現れると、生成された観測からそれを取り込み、生成世界とともに相互作用可能な範囲を広げる。一方、以前に引き起こした状態変化は動画区間をまたいで引き継がれ、相互作用の結果がその後の世界の発展に持続的に組み込まれる。更新された世界状態を、カメラの行動軌跡に沿って粗い条件付け動画へ描画し、動画生成器が、状態に表現されていない外観、動き、相互作用の細部を補う。 実験は、Oneiraが新たに生成した物体との直接的で一貫した相互作用を可能にすると同時に、長い時間範囲にわたって以前の相互作用の効果を保つことを示している。プロジェクトページ:https://madaoer.github.io/projects/oneira

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended generation does not imply full interaction: as a generated world expands, newly created content through navigation should expand what the agent can act upon, and as the agent changes the world, those changes should become persistent parts of the environment rather than transient visual effects. We characterize these two requirements as Open-World Interactivity, where newly generated or encountered entities are incorporated into the actionable world, and Persistent State, where interaction outcomes are committed to the world state and continue to influence subsequent observations and interactions. We present Oneira, an interactive video world model that closes the loop between generation and interaction through an explicit, extensible world state managed by a coding agent. Given the current observation and an action or high-level goal, the agent reads the world state, grounds the relevant entities, plans the interaction, and writes its outcome back into a world state table. When exploration reveals new objects, the agent incorporates them from generated observations, allowing the interaction space to expand with the generated world. Meanwhile, previously induced state changes are carried across video segments, making the consequences of interaction persistent parts of subsequent world evolution. The updated world state is rendered along the camera action trajectory into a coarse conditioning video, from which a video generator fills in the appearance, motion, and interaction details not represented in the state. Experiments show that Oneira enables direct and consistent interaction with newly generated objects, while preserving the effects of prior interactions over long horizons. Project page: https://madaoer.github.io/projects/oneira

著者のコメント

Project page: https://madaoer.github.io/projects/oneira

arXiv ID: 2610.01614 / 要約の誤りについて