arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

ロボットを動かさず人の操作でVLAを追加学習

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Zimu Han, Yiming Zeng, Jiyao Zhang, Zihao Zhao, Yuanfei Wang, Yixiang Jin, Shiqi Li, Shuangben Chen, Wei Huang, Ruodai Li, Hui Shen, Hao Dong

この論文をやさしく読む

ひとことで言うと

人が手持ちの操作器具を動かしながら、ロボットを実行せずに現在の方策が苦手な場面を収集します。

何に役立つ?

特定の作業へVLAを適応させる際、実機を繰り返し動かすデータ収集の負担を減らす用途が考えられます。

この研究の面白いところ

人の軌跡と方策の推論の差で収集場面を選び、進捗評価器の改善を別のフィードバックで行います。

どこまで分かった?

実機四タスクで改善し、一つの片付けタスクではHG-DAggerと比較しています。「ロボット不要」は収集方式の説明で、評価自体は実機を含みます。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模な視覚・言語・行動モデル(VLA)はロボット操作に強力な事前知識を与えるが、特定の運用環境への適応にはなお課題がある。タスク固有の実演による教師ありファインチューニング(SFT)は運用に向けた一歩となるものの、2つの継続的な制約がある。静的データは分布外の状態を十分に網羅できず、標準的な模倣学習の目的関数は、進展につながる行動と有用性の低いデータを区別しない。対話的な追加学習はこの制約に対応できるが、通常は実機ロボットでの方策の反復実行と人の介入を必要とする。 本研究では、ロボットを使わずに人を学習ループに組み込むVLA追加学習のため、方策に誘導される汎用操作インターフェース(UMI)の枠組みHIL-UMIを提案する。手持ちのUMIによる実演中、同じ観測ストリームを現在の方策に入力するが、その予測は実行しない。Energy Scoreが人の行動軌跡と方策の推論結果を比較し、両者の食い違いが分布外領域を示すとデータ収集を開始する。別のフィードバックループでは、オンラインのアドバンテージ予測が低い箇所を使って、進展に基づくアドバンテージ推定器の改善に重要な区間を特定する。更新した推定器は、基本の実演と新たな方策データをバランスよく混ぜたデータを使い、アドバンテージ条件付き行動クローニングを導く。この設計は、人を学習ループに組み込む手法の反復性と方策を考慮する性質を保ちながら、データ収集をロボットの運用から切り離す。 長い手順や精密な操作を含む4つの実世界タスクでの実験では、HIL-UMIはSFTに対して一貫した改善を示し、対象を絞った収集とアドバンテージ推定の改善の両方が有効だった。さらに、Clean Up TableではHG-DAggerを上回り、1フレーム当たりの収集時間も短かった。この結果は、操作者や場所をまたいでVLA追加学習を拡張できる可能性を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limited coverage of out-of-distribution states, and standard imitation objectives do not distinguish progressing behavior from less useful data. Interactive post-training can address these limitations, but typically requires repeated policy execution and human intervention on a physical robot. We introduce HIL-UMI, a policy-guided Universal Manipulation Interface (UMI) framework for robot-free human-in-the-loop VLA post-training. During handheld UMI demonstrations, HIL-UMI queries the current policy on the same observation stream without executing its predictions. The Energy Score compares the human action trajectory with policy inference and triggers collection when their discrepancy indicates an out-of-distribution region. In a separate feedback loop, low online advantage predictions identify essential segments for refining a progress-based advantage estimator. The updated estimator then guides advantage-conditioned behavioral cloning using a balanced mixture of base demonstrations and new policy data. This design preserves the iterative and policy-aware nature of human-in-the-loop learning while decoupling data collection from robot deployment. Experiments on four real-world tasks spanning long-horizon and precise manipulation show that HIL-UMI achieves consistent improvement over SFT and benefits from both targeted collection and advantage refinement. Moreover, HIL-UMI outperforms HG-DAgger on Clean Up Table with lower per-frame collection time, suggesting a scalable path for VLA post-training across operators and locations.

arXiv ID: 2609.20659 / 要約の誤りについて