arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

世界モデルの状態の一貫性と操作への応答を評価

HappyWorld-Bench

Zhiqi Bai, Junai Cai, Yixin Chen, Jingrun Du, Tao Feng, Wei Gong, Siyuan Huang, Xiao Lin, Jiaheng Liu, Jun Luo, Yongzhe Lyu, Liya Ma, Zenan Meng, Lin Qu, Wenbo Su, Jiaming Wang, Qinghe Wang, Shaofei Wang, Yanghai Wang, Zequn Wang, Ziming Wang, Hu Wei, Jiangtao Wu, Ruiqi Wu, Jiaxin Xie, Yuchi Xu, Ze Xu, Chengting Yu, Liangyu Yuan, Gang Zeng, Yawen Zeng, Xingyao Zhang, Zizheng Zhang, Bo Zheng, Jiancheng Zhu, Song-Chun Zhu

この論文をやさしく読む

ひとことで言うと

見栄えのよい世界を作れるだけでなく、動き回ったり変更したりしても矛盾しないかを測る評価集です。映像、空間、行動を伴うモデルを分けて試しています。

何に役立つ?

世界モデルを比較するとき、見た目だけでは分からない状態保持や操作への反応の問題を見つけるために役立ちます。人の比較と自動評価を併用しています。

この研究の面白いところ

長く探索する、元の場所へ戻る、ルールを変えるといった操作を評価へ入れています。生成直後の品質と、使い続けたときの信頼性を分けて測ります。

どこまで分かった?

報告値はこの評価集と選んだモデル群での結果です。空間モデルの配置精度と編集実行率は別指標で、映像や身体性モデルの総合正答率ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

世界モデルの評価には、生成する世界の品質に加えて、探索、相互作用、変更の下での一貫性と応答性を評価する必要がある。本研究は、エージェントが相互作用する際にも生成世界が信頼できる状態を保つかを測る、包括的なベンチマークHappyWorld-Benchを導入する。設計は、生成による構築から統一的な世界モデリングまでの6つの世界能力(W1~W6)の階層的枠組みに基づき、映像世界モデル、空間世界モデル、身体性を持つ世界モデルという独立した3つの評価系列で具体化する。 HappyWorld-Benchは、1138の映像プロンプト、300の空間場面、254の身体性を持つテストケースからなる。3系列すべてでHappyWorld-Arenaを構築・運用し、人によるA/B比較を整理してモデル単位のElo評価を求める。これらは、行動の正しさを捉える新しい自動指標を補完する。この統一枠組みで、14の映像世界モデル、9つの空間システム、8つの身体性を持つ候補を評価する。 結果は3系列すべてに残る信頼性の不足を明らかにする。映像モデルは長い生成系列や再訪で一貫性が下がり、空間モデルの配置精度は最高でも70.14%、編集実行は73.33%である。身体性を持つモデルは、多段階の行動にわたる状態の保持や、変更された行動条件・物理法則への正確な応答に苦戦する。これらの知見は、視覚品質だけでなく、状態の一貫性と、行動・介入への応答の正しさによって世界モデルを評価する必要を示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, instantiated across three independent evaluation tracks: video world models, spatial world models, and embodied world models. HappyWorld-Bench comprises 1,138 video prompts, 300 spatial scenes, and 254 embodied test cases. Across all three tracks, we build and operate HappyWorld-Arena to organize human A/B comparisons and derive model-level Elo ratings, which complement newly designed automated metrics that capture behavioral correctness. We evaluate 14 video world models, 9 spatial systems, and 8 embodied candidates under this unified framework. Results reveal remaining reliability gaps across all three tracks: video models exhibit reduced consistency during extended rollouts and revisits, spatial models achieve at best 70.14% placement accuracy and 73.33% edit execution, and embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physical rules. These findings highlight the need to evaluate world models not only by visual quality, but also by state consistency and the correctness of their responses to actions and interventions.

arXiv ID: 2609.24308 / 要約の誤りについて