手先の3次元形状変化で異なる身体間のロボット学習をつなぐ
GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments
この論文をやさしく読む
ひとことで言うと
人の手とロボットの手など、形が違う手先の動きを共通の表現にして、動画からロボットの行動を学びやすくする手法です。
何に役立つ?
行動ラベルが付いていない人間の一人称動画も含め、異なる身体のデータをVLAの事前学習に活用するための方法になります。下流評価と実環境評価の成功率が要旨に報告されています。
この研究の面白いところ
画像で捉える場面全体の変化と、3次元形状で捉える指先の細かな動きを組み合わせます。点群を加えるだけでは身体間で意味がそろわない問題に、統一運動表現で対処しています。
どこまで分かった?
68.3%はRoboCasa-GR1、75.5%は実環境での別々の評価値です。要旨には試行数やタスクの内訳、未知の手や環境でのすべての成功率はなく、任意の身体へそのまま汎化するとの意味ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
複数の身体構成を含むデータセットから大規模な視覚・言語・行動(VLA)モデルを学習することは、エンドエフェクタごとに行動空間が異なるため、依然として難しい。潜在行動モデル(LAM)は多様な動画データから身体構成に依存しない行動表現を学習できるが、既存の画像ベースのLAMでは、エンドエフェクタの細かな関節運動、特に人間の手や器用なロボットハンドの指レベルの幾何学的変化を十分に捉えられないことが多い。 この制約に対処するため、画像ベースの潜在行動にエンドエフェクタの3次元幾何学的運動を加える、幾何を考慮した潜在行動モデリングの枠組みGALAを提案する。しかし、点群を単純に取り入れるだけでは、細かな行動表現は得られても共通する意味が乏しく、異なる身体構成を横断した事前学習が妨げられる。そこで、細かな運動情報を保持しながら潜在行動の身体構成間の汎化能力を改善する、統一エンドエフェクタ運動表現(UEMR)を導入する。 UEMRを基盤として、GALAは場面全体のダイナミクスを捉える視覚的潜在行動と、共通する細かなエンドエフェクタの関節運動を捉える幾何学的潜在行動を組み合わせる。これにより、行動情報を持たない人間の一人称視点動画を含む、複数の身体構成のデータからVLAを事前学習するための有効な教師信号を提供する。細かな運動のプロービング、異なる身体構成間の検索、および下流のVLA評価の実験は、身体構成を横断して汎化可能な細かな運動をモデル化するGALAの有効性を示し、RoboCasa-GR1で成功率68.3%、実環境で成功率75.5%を達成する。コード、付録、デモはhttps://puzhenyuan.github.io/GALA-website/で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA's effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate. Code, appendix, and demos are available at https://puzhenyuan.github.io/GALA-website/.
arXiv ID: 2609.21948 / 要約の誤りについて