arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

未来の動画を作らずロボットの動きを予測するMoWAM

MoWAM: Explicit Future Motion Prediction for Efficient World Action Models

Jiayu Wang, Bin Zhu, Yue Yu, and Jingjing Chen

この論文をやさしく読む

ひとことで言うと

ロボットが未来の動画全体を生成する代わりに、未来の自分の動きを明示的に予測します。

何に役立つ?

考えられる用途は、将来予測を保ちながらロボット操作モデルの推論負担を抑えることです。

この研究の面白いところ

小さな運動表現を使って複数の行動候補を作り、作業進捗を見積もる検証器で選べる点が特徴です。

どこまで分かった?

LIBERO系と実機タスクで頑健性や平均成功率の改善を報告しています。具体的な推論時間や成功率の数値は要旨には示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

世界行動モデル(WAM)は、将来の動力学を取り入れてロボットの方策学習を改善するが、推論時に未来の動画を明示的に生成すると、計算負荷が大きくなる。未来の生成を取り除けば効率は上がるものの、将来の動力学は観測特徴に暗黙的に符号化されるだけとなり、分布の変化に対する頑健性が制限される可能性がある。 未来の動画生成を、将来の動作の明示的予測に置き換える効率的なWAM、MoWAMを提案する。未来の場面全体を再構成する代わりに、構造化されたロボットの動作を未来のコンパクトな抽象表現としてモデル化し、現在の場面と相互作用の制約の下で、ロボットがどう変化すると予想されるかを捉える。Mixture-of-Transformer構造が、学習時に将来の視覚的動力学を学びながら動作と行動を同時に予測することで、未来の明示的表現を残しつつ、推論時の動画生成を完全に取り除ける。 コンパクトな動作表現により、動作と行動の対を複数サンプリングし、動作を考慮する課題進捗検証器で候補を選ぶ、効率的な推論時の計算量拡大も可能になる。LIBERO、LIBERO-Plus、実環境の操作課題での実験では、MoWAMは分布内で高い性能を達成し、分布外での頑健性を改善し、代表的なWAM比較手法より高い実環境の平均成功率を示す。さらに、探索する候補を増やすと性能も改善し、明示的な未来動作が、推論時の計算量拡大に効果的かつ効率的な基盤を与えることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet explicitly generating future videos at inference introduces substantial computational overhead. Removing future generation improves efficiency, but leaves future dynamics only implicitly encoded in observation features, which can limit robustness under distribution shifts. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Instead of reconstructing the complete future scene, MoWAM models structured robot motion as a compact abstraction of the future, capturing how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing video generation to be removed entirely at inference while retaining an explicit representation of the future. The compact motion representation further enables efficient inference-time scaling by sampling multiple candidates of motion and action pairs and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines. In addition, performance improves as more candidates are explored, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.

arXiv ID: 2609.20709 / 要約の誤りについて