物体操作に伴う形状変化を学ぶロボット方策
HABILIS Brain 0: Geometry-Change Supervision for Vision-Language-Action and Residual Flow Recovery
この論文をやさしく読む
ひとことで言うと
ロボットが物体を操作した後の形状変化を予測し、それを行動の学習に使う方法を提案した。
何に役立つ?
ロボットや一人称視点の映像を使った事前学習から、物体操作の方策へつなぐ設計の参考になる。LIBEROで成功率を評価した。
この研究の面白いところ
将来の画像は学習目標の作成にだけ使い、実際の推論では現在の観測で動く。追加の残差速度方策を組み合わせると、LIBERO成功率が95.20%から99.55%になった。
どこまで分かった?
報告された成功率はLIBERO上の結果である。実機や別の課題での性能は要旨には記載されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚・言語・行動の方策には幾何学的な教師信号が役立つが、現在のフレームの形状だけでは、物体操作に伴って何が変わるかを明示できない。この設計は、ロボット固有の行動と対応付ける前に、ロボット映像と一人称視点の映像をまたいで事前学習できる、機体に依存しない視覚表現を学ぶことを目指している。 著者らは、現在の観測から、複数視点で見た将来と現在の幾何学的変化を表すトークンを予測するGeometry-Change VLA(GC-VLA)を導入する。オフラインのフレーム対から、標準的な0.5秒先の予測区間を定める。将来の観測は学習目標の作成だけに使う。第1段階では幾何学的変化の視覚言語モデル(GC-VLM)を学習する。第2段階では連続的なActionExpertを導入してロボットの行動と対応付け、この時点では行動フローの勾配をVLMとの境界で止める。第3段階では、この勾配で学習可能なVLM部分もActionExpertと共同で更新する。第4段階ではGC-VLAを固定し、閉ループのフィードバックから学んだ二値の介入ルーターと、範囲を制限した単一の残差速度方策を用いるGeometry-Conditioned Residual Flow(GCRF)を適用する。 GC-VLAはLIBEROで95.20%の成功率を達成し、GCRFを組み合わせると99.55%になった。推論時には現在の観測と学習済みのGC表現を使い、オフラインの目標作成用エンコーダーは実行しない。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-language-action policies benefit from geometric supervision, but current-frame geometry alone does not explicitly describe the changes associated with manipulation. This design is motivated by the goal of learning an embodiment-agnostic visual interface that can be pretrained across robot and egocentric video before robot-specific action alignment. We introduce Geometry-Change VLA (GC-VLA), which learns to predict multiview future-current geometry-change tokens from current observations. Offline frame pairs define a nominal 0.5-second prediction horizon; future observations are used only to construct training targets. Stage 1 trains a geometry-change vision-language model (GC-VLM). Stage 2 introduces a continuous ActionExpert and aligns it with robot actions while stopping action-flow gradients at the VLM interface. Stage 3 enables these gradients to update the trainable VLM components jointly with the ActionExpert. Stage 4 freezes GC-VLA and applies Geometry-Conditioned Residual Flow (GCRF), using a binary intervention router and a single bounded residual velocity policy learned from closed-loop feedback. GC-VLA achieves 95.20% success on LIBERO, and GC-VLA with GCRF achieves 99.55%. Inference uses current observations and the learned GC representation without executing the offline target encoders.
arXiv ID: 2609.25558 / 要約の誤りについて