arXiv論文メモ
新着一覧
cs.AI / cs.CV · 査読状況未確認

物体が見えなくなっても存在すると判断する世界モデルを学習

Training Object Permanence in World Models

Haotian Zhang, Fengyuan Yu, Dezhi Luo, Haoran Sun, Zehong Zhao, Qingying Gao, Yihan Li, Siyuan An, Huayi Qin, Yilan Zhang, Zhengze Jiang, Pinyuan Feng, Renrui Zhang, Ziyu Guo, Letian Wang, Mengyue Yang, Kangfu Mei, Maijunxian Wang, Ran Ji, Vikash Kumar, Freda Shi, Chandra Sripada, Vincent C. Muller, Philip Torr, Alan Yuille, Nikolaus Kriegeskorte, Felix Juefei-Xu, Lvmin Zhang, Jieneng Chen, Yilun Du, Hokin Deng

この論文をやさしく読む

ひとことで言うと

動画生成モデルに、隠れた物体も存在し続けるという理解を学ばせるための課題とデータを作った研究。

何に役立つ?

動画モデルの物理的な推論を、見た目の自然さとは別に評価・学習するためのデータ基盤になる。

この研究の面白いところ

150課題から150万件の学習データを作り、14モデルを300問の試験と人による一対比較で評価した。

どこまで分かった?

要旨に示される順位はこの試験と一対比較に基づく。あらゆる物理現象への一般化を示す結果は記載されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

物体の永続性と固体性は、人間が持つ認知上の前提を特徴づける。最近の研究では、現在の世界モデルの代表例である動画生成モデルに推論能力が現れ始めたと報告されており、人間に似た物理的知能を築く候補となっている。動画モデルには物体の永続性が自発的に備わっているのだろうか。ない場合、基礎的な認知に着想を得たデータで学習できるのだろうか。本研究は、認知科学に着想を得て手作業で設計した150の課題を六つの認知カテゴリに分けた、WROP(World Reasoning with Object Permanence)というデータ基盤を導入する。各課題の認知的な構造を保ちながら速度、照明、カメラ角度など本質的でないパラメーターをランダム化するBlenderの生成器を作り、課題ごとに1万件以上のサンプルを得る。150万件の学習コーパスと300問の試験を公開する。この試験で、参照画像から動画を作るモデル3種、編集モデル7種、続きの動画を作るモデル4種の計14モデルを評価した。そこには本研究の160億パラメーターの世界モデルPWM-WROPも含まれる。モデル名を伏せた一対比較のElo評価では、PWM-WROPは動画継続モデルの中で1位、全体では3位だった。上位には、統計的に同順位の参照画像から動画を作る2モデルだけがあった。データ、試験、モデルの回答、得点、重み、AWS Trainium2上の純粋なPyTorchによる学習環境PWMを公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Object permanence and solidity are hallmarks of human cognitive priors. Recent studies show that video generation models, a paradigmatic class of current world models, have begun to show emerged reasoning abilities, making them ideal candidates for building human-like physical intelligence. Do video models have emerged object permanence in them? If not, could we train them with a core-cognition inspired dataset? We introduce WROP (World Reasoning with Object Permanence), a data infrastructure of 150 hand-designed cognitive science inspired tasks, divided into six cognitive categories. We build Blender generators that randomize speed, lighting, camera angle, and other nuisance parameters while preserving each task's cognitive structure, yielding 10,000+ samples per task. We release a 1.5M-sample training corpus and a 300-question exam. On this exam we evaluate 14 video models: 3 reference-to-video, 7 edit, and 4 continuation, among which PWM-WROP, our 16B world model. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind only a statistical tie between two reference-to-video models. We release the data, exam, model answers, scores, weights, and PWM, our native-PyTorch training stack on AWS Trainium2.

著者のコメント

26 pages, 9 figures, 5 tables. Project page: https://object-permanence.world

arXiv ID: 2609.28654 / 要約の誤りについて