物体の3次元位置を基盤にしたロボット操作モデル
Grounded Action Model: 3D Grounding as a Foundation for Robotics
この論文をやさしく読む
ひとことで言うと
操作対象の見た目だけでなく、3次元の位置や形状を明示的に取り込んでロボットの動作を決める研究です。言葉・点・枠による指定を共通の物体表現に変換します。
何に役立つ?
対象の置き場所や見た目が変わる操作を評価する際に役立ちます。上位の計画器と組み合わせた長い作業への利用も検討され、Frankaでステップ完了率が報告されています。
この研究の面白いところ
行動方策を乱れのないシーンだけで学習させても、ランダム化したシーンで性能を保つ点を調べています。対象の指定方法を統一することで、自律操作と計画器からの指示を同じ仕組みで扱います。
どこまで分かった?
ベンチマークの成功率と、実機の試行成功数やステップ完了率は異なる指標です。Frankaの49.8%は分布外での各ステップの完了率であり、長い作業全体の成功率ではありません。要旨は任意の物体や環境への保証を示していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
物体操作の方策には、どの物体が重要で、それがどこにあるかを把握する能力が必要である。しかし、現在のロボット基盤モデルが利用する事前学習済みの基盤、すなわち視覚・言語・行動モデル(VLA)における言語から、世界・行動モデル(WAM)における動画生成までの各種基盤は、実空間の尺度に基づくこの対応付けを直接には要求せず、その獲得をロボットの実演からの暗黙的な学習に委ねている。本研究は、3次元の対応付けを基盤とする新しいロボット基盤モデルの枠組み、Grounded Action Models(GAM)を提案する。GAMには言語、点、またはボックスのプロンプトで条件を与えられ、それらはまず、選択された物体を中心とする共通表現へ変換される。この表現は対象に焦点を当てた視覚特徴と、実空間の尺度を持つ物体の幾何情報を捉え、複数ストリームのTransformerを通じてロボット状態の履歴と統合され、一連の行動をまとめて予測する。GAMは自律実行できるだけでなく、上位の計画器が複数の入力形式を通じて制御する下位コントローラーとしても機能し、長い時間範囲にわたる操作や記憶を必要とする操作を可能にする。 RoboTwin 2.0では、GAMは50タスク平均で55.3%の成功率を達成した。Spatial Forcingは52.0%である。行動方策は乱れのないシーンの実演だけで学習しているが、シーンをランダム化した条件でも47.6%を達成し、Abot-M0の30.4%を上回った。LIBERO-PROでは、16種類の摂動設定を通じて平均成功率61%という最高水準の結果を得た。π₀.₅は53%であり、対象を移動した場合や新たに指定した場合に改善が最も大きかった。2台の実ロボットでも評価した。双腕YAMでは視覚的な変化の下で20回中17回の成功を維持し、π₀.₅は20回中4回だった。一方、Franka上でMolmo2計画器と組み合わせた場合、長期的な操作や記憶に依存するタスクで、分布内(ID)のステップ完了率64.7%、分布外(OOD)では49.8%を達成した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects. This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs. 30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for $\pi_{0.5}$) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for $\pi_{0.5}$, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.
arXiv ID: 2609.23863 / 要約の誤りについて