ロボットの動作も画像パッチとして生成するPatchWAM
An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM
この論文をやさしく読む
ひとことで言うと
ロボットの動作を画像の小領域と同じ形式にして、視覚生成モデル一つで動作と次の場面を予測する方法。
何に役立つ?
視覚モデルをロボット制御へ拡張する際、専用の動作モデルを増やさない設計を検討する材料となる。
この研究の面白いところ
動作を固定写像でパッチ化するだけで画像予測と動作生成を同じ経路に載せ、複数のロボット課題で高い成功率を報告した。
どこまで分かった?
91.8%と96.12%は全データと追加の拡張実演を使った設定でのベンチマーク結果。別条件への一般化や実機性能は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
生成型の視覚モデルは物理的な動きの表現を学ぶ基盤になるが、連続的な制御に広げるとき、画像予測と動作生成には別々の計算経路が必要なのかという問題が生じる。従来の方法は、低次元の状態と高次元の視覚表現をつなぐため、学習可能な動作出力部や別の動作専門モデルを設けることが多い。本研究は、動作を適合する表現で示せば、既存の視覚モデルの能力を制御にも使えるかを調べる。提案するPatchWAMは、Action-as-Patchという固定した写像を使い、連続動作を別種の画像パッチとして扱う。これにより一つのモデルが、ロボットの次の動きと、その後に場面がどう見えるかをともに予測できる。専用の動作出力部や別の動作専門モデルなしに、画像予測と動作生成を同じ生成過程に組み込む。学習窓を間引いた実験では条件をそろえた二専門モデルの制御方式を上回り、全データと追加の拡張実演を使ったベンチマークではLIBERO-Plusで91.8%、RoboTwin 2.0で96.12%の成功率を得た。結果は、生成モデルを新しい信号へ拡張する制約が、モデル容量だけではなく、その信号をどのような接続表現で与えるかにあることを示唆する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Generative visual models offer a foundation for learning representations of physical dynamics, yet their extension to continuous control raises a fundamental question: do visual prediction and action generation require separate computational pathways? Existing approaches usually introduce trainable action heads or separate action experts to bridge low-dimensional states and high-dimensional visual representations. In this work, we explore whether the visual backbone's existing capacity can also support control when actions are expressed in a compatible representation. Thus, we introduce PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed mapping called Action-as-Patch. This allows a single model to predict both how the robot should move and what the scene may look like afterward. Visual prediction and action generation become parts of the same generative process, without a dedicated action head or separate action expert. Experiments with subsampled training windows show gains over a matched dual-expert control, while benchmark evaluations reach 91.8% success rate on LIBERO-Plus and 96.12% on RoboTwin 2.0 in a full-data setting with additional augmented demonstrations. More broadly, the result suggests that capability need not be added where it can be inherited: the constraint on extending a generative backbone is the interface a new signal is written in, not the capacity to model it.
arXiv ID: 2609.25961 / 要約の誤りについて