要所の動作判断と軌道生成を分けたロボット操作モデル
H-VLA: Hierarchical Vision-Language-Action Model with Key-Action Reasoning and Motion Planning in a Unified Action Space
この論文をやさしく読む
ひとことで言うと
文章と画像からロボットを動かす際、次の重要な動作を決める段階と細かな動きを作る段階を分けた研究。
何に役立つ?
物体の位置や視点が変わるロボット操作で、意味的な判断と軌道生成を両立する方法の検討に役立つ。
この研究の面白いところ
重要動作と細かな運動を別モデルで扱いながら、カメラ中心の共通行動空間で複数の条件をつないでいる。
どこまで分かった?
報告値はSimplerEnvとAgilex実機の指定課題での結果であり、あらゆるロボットや環境への汎化を示すものではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚・言語・行動(VLA)モデルはロボット操作で有望だが、既存手法の多くは文章指示と画像観測から細かな動作へ直接写像する。この方式では、主に画像と言語の理解のために学習された視覚言語モデル(VLM)の意味的な推論能力が弱まり得る。また、物体位置、場面の配置、ロボットの形、カメラ視点が変わると不安定になり得る。本研究は、高い段階での重要動作の推論と、低い段階での運動生成を分ける階層型VLAの枠組み H-VLA を提案する。 H-VLAは、次の操作上の小目標として重要動作を予測するKey-Action Model、その予測を条件に細かな将来動作を生成するMotion Planning Model、データセット・ロボット本体・視点をまたいで一貫した表現を与えるカメラ中心の統一行動空間を組み合わせる。事前学習では重要動作の推論を、微調整では細かな運動生成を重視する2段階の学習を採用する。 実験では、SimplerEnvのGoogle Robot視覚一致で91%、Google Robotの変種をまとめた評価で84%、WidowX視覚一致で81%に達した。Agilexの実機ロボット課題では、評価条件と同じ分布、分布外の物体位置、分布外の場面・物体という3条件で、最も強い比較手法をそれぞれ10、47、16ポイント上回った。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but many existing methods still rely on direct mappings from language and visual observations to dense actions. This formulation can weaken the semantic reasoning capability inherited from pre-trained Vision-Language Models (VLMs), which are mainly optimized for visual-linguistic understanding rather than low-level control, and becomes fragile under spatial variations, including changes in object positions, scene layouts, robot embodiments, and camera viewpoints. To address these limitations, we propose H-VLA, a hierarchical VLA framework that decouples high-level key-action reasoning from low-level motion generation. H-VLA combines a Key-Action Model for predicting a key-action as the next manipulation subgoal, a Motion Planning Model for generating dense future actions conditioned on the predicted key-action, and a Unified Camera-Centric Action Space for consistent representation across datasets, embodiments, and viewpoints. We further adopt a two-stage training strategy that emphasizes key-action reasoning during pre-training and dense motion generation during fine-tuning. Experiments show that H-VLA achieves strong performance on SimplerEnv, reaching 91% on Google Robot visual matching, 84% on Google Robot variant aggregation, and 81% on WidowX visual matching. On Agilex real-robot tasks, H-VLA improves over the strongest baseline by 10, 47, and 16 percentage points under in-distribution, out-of-distribution position, and out-of-distribution scene/object settings, respectively.
arXiv ID: 2609.22895 / 要約の誤りについて