ロボットの世界・行動モデルの設計要素を分けて比較する
What Matters in Designing World Action Models: An Empirical Study
この論文をやさしく読む
ひとことで言うと
ロボットが環境の変化を予測し行動を作るモデルで、構造、内部表現、学習目標のどれが性能に影響するかを分けて調べる研究です。
何に役立つ?
考えられる用途は、WAMを設計するときの比較条件の整理や設計選択です。複数の変更をまとめて評価すると分かりにくい寄与を切り分けます。
この研究の面白いところ
六つの因果構造、八つの潜在表現、四つの学習目的という別々の設計軸を統制して比較しています。
どこまで分かった?
要旨には、どの選択が最良だったかや性能差の数値がありません。DROIDによる確認は実ロボットのデータを使った検証であり、新たな実機運用実験が行われたとは書かれていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
World Action Models(WAM)は、汎化可能なロボット制御の有望な枠組みとして登場している。WAMシステムは増えているが、既存研究は構造や学習戦略など複数の設計上の選択を一つにまとめたシステムを提示することが多く、それぞれの寄与を切り分け、別の設計と体系的に比較することが難しい。本研究では、こうした設計上の選択を分離した統制研究を提示し、実験上の効果だけでなく、それらがWAMをどのように、なぜ形作るかも分析する。 具体的には、WAM構築における三つの基本的な問いに焦点を当てる。(1)世界モデリングと行動生成の相互作用をどのような因果構造に従わせるべきか、(2)どの潜在空間で世界モデリングを行うべきか、(3)世界と行動をモデル化する目的関数の違いが、モデルの振る舞いと性能にどう影響するか、である。代表的な三つのベンチマークRoboCasa-GR1、LIBERO、LIBERO-Plusで構造を統制した実験を行い、既存WAMで一般的な設計を含む六つの因果構造、八つの潜在表現、四つの学習目的を体系的に比較する。さらに、DROIDデータセットの実ロボットデータで主要な知見を検証する。中心的な設計の選択が世界・行動モデリングにどう影響するか、今後のWAMシステム開発をどの原則で導けるかについて、体系的な理解を提供することを目指す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
World Action Models (WAMs) have emerged as a promising paradigm for generalizable robot control. Despite the growing number of WAM systems, existing works often introduce unified systems that bundle together multiple design choices, such as architecture and training strategy, making it difficult to isolate individual contributions and systematically compare alternative designs. In this work, we present a controlled study that disentangles these design choices and analyzes not only their empirical effects, but also how and why they shape WAMs. More specifically, we focus on three fundamental questions in building WAMs: (1) what causal structure should govern the interaction between world modeling and action generation? (2) in which latent space should world modeling be performed? and (3) how do different world-action modeling objectives affect model behavior and performance? Through structurally controlled experiments on three representative benchmarks, RoboCasa-GR1, LIBERO, and LIBERO-Plus, we systematically compare six causal structures, eight latent representations, and four training objectives, covering popular design choices in existing WAMs. We further validate our key findings on real-robot data from the DROID dataset. We hope to provide a systematic understanding of how core design choices affect world-action modeling and what principles can guide the development of future WAM systems.
arXiv ID: 2609.24048 / 要約の誤りについて