arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

高自由度の器用な操作へ向けたVLA追加学習

Towards High-DoF Dexterous Manipulation through VLA Post-Training

Junlei Zhu, Shenzhe Yao, Chaogui Huang, Wenkai Zhu, Jingwei Peng, Guanqi He, Soren Schwertfeger, Jiahao Chen and Yide Liu

この論文をやさしく読む

ひとことで言うと

多くの関節を持つロボットハンドへ、既存の視覚言語行動モデルを適応させる4段階の追加学習手順です。

何に役立つ?

両手での受け渡し、手内での向き変更、道具使用といった実機操作の調整に役立つ方法です。

この研究の面白いところ

手の協調動作を表す圧縮表現を学び、人が操作を引き継ぐ際の姿勢のずれを滑らかにつなぎ、その表現内で残差強化学習を行います。

どこまで分かった?

5つの実機課題で各20試行すべて成功したと報告しています。指定された追加学習予算と評価試行での成績であり、任意の作業や環境での100%成功を保証するものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

模倣学習で得られた視覚・言語・行動(VLA)基盤モデルは、複数の課題やロボット形態にまたがるデータを大規模化することで、幅広い操作能力を獲得する。しかし、特定の下流課題とハードウェアへ信頼性高く展開するには、なお追加学習が必要である。器用なハンドでは、広い行動レパートリーと高い自由度によって、広大で構造化された行動空間が生じるため、この適応は特に難しい。中心的な障害は三つある。オープンソースのVLAは高自由度ハンド用の行動インターフェースを本来備えていないこと、人間が介入するDAggerの引き継ぎ時にジェスチャーの不一致が命令の不連続性を生み、修正軌道を汚染すること、そして生の関節空間での強化学習はサンプル効率が低いことである。 本研究では、学習した時間的ハンド行動コーデック、教師あり微調整、DAgger、実世界での残差強化学習からなる、統一的な4段階の追加学習パイプラインを提案する。このコーデックは、事前学習済みVLAを器用なハンドへの絶対指令に適応させる。バッファ付きロールバック、姿勢合わせ、滑らかな命令ブレンディングによって、連続的で課題に関連したDAgger修正を可能にする。また潜在残差RLは、コーデックが捉えた協調的なハンド運動に探索を限定する。 両手での移送、手中での再配向、道具使用にまたがる多様な実世界の5課題で、このパイプラインを評価した。報告された追加学習予算の範囲では、得られた方策は各課題20試行で、評価したすべての課題において成功率100%を達成した。これらの結果は、VLA基盤モデルを信頼性の高い実世界の器用な操作へ適応させる実用的な道筋を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Imitation-learned vision--language--action (VLA) foundation models acquire broad manipulation capabilities by scaling robot data across tasks and embodiments, but reliable deployment on a specific downstream task and hardware platform still requires post-training. Dexterous hands make this adaptation particularly difficult: their broad behavioural repertoire and high degree of freedom create a large and structured action space. Three obstacles are central: open-source VLAs do not natively provide an action interface for high-DoF hands; gesture mismatch during human-gated DAgger takeover creates command discontinuities and contaminates corrective trajectories; and reinforcement learning in the raw joint space is sample-inefficient. We present a unified four-step post-training pipeline comprising a learned temporal hand-action codec, supervised fine-tuning, DAgger, and real-world residual reinforcement learning. The codec adapts a pretrained VLA to absolute dexterous-hand commands. Buffered rollback, pose alignment, and smooth command blending enable continuous, task-relevant DAgger corrections, while latent residual RL confines exploration to coordinated hand motions captured by the codec. We evaluate the pipeline on five diverse real-world tasks spanning bimanual transfer, in-hand reorientation, and tool use. Within the reported post-training budgets, the resulting policies achieve 100\% success on every evaluated task over 20 trials per task. These results provide a practical path for adapting VLA foundation models to reliable real-world dexterous manipulation.

著者のコメント

29pages, 10 figures

arXiv ID: 2609.19666 / 要約の誤りについて