arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

人の修正の確かさに合わせて学ぶロボット方策

BEE: Intervention-Adaptive Real-World Reinforcement Learning with Vision-Language-Action Models

Weihui Zhao, Xiaohan Yan, Zunian Wan, Xuan Du, Zhaozhan Chi, Jianbo Mao, Ruipu Wu, Rushuai Yang, Houlin Li, Shukai Yang, Jing Wu, Yuxiang Yan, Yongcheng Liu, Chuankang Li, Guanghui Ren, Wei Shan, Maoqing Yao

この論文をやさしく読む

ひとことで言うと

ロボットの動作を人が直したとき、その修正がどの方向で一貫しているかを見て、従う強さを調整する学習方法。

何に役立つ?

実機での自由な試行を増やしにくい物体操作で、人の修正を効率よく使う方策改善に役立つ可能性がある。報告された成功率は、評価した4課題での平均である。

この研究の面白いところ

人の修正をそのまま真似せず、動作の次元ごとに制約の強さを変える。条件をそろえた4課題で平均成功率91.2%、比較手法は57.5%と42.1%だった。

どこまで分かった?

実世界3課題とシミュレーション1課題での評価であり、別の作業に同じ改善が出るかは要旨からは分からない。人による修正を得る過程自体は必要である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚・言語・動作(VLA)モデルは長い工程の物体操作を扱えるが、成功は、数ミリメートルの誤差がそれまでの進捗を台無しにするような、精度が重要な少数の段階に左右される。オンライン強化学習はまさにそうした動作を改善できるが、実際のロボットで自由に探索する費用は高すぎるため、人による修正が欠かせない。しかし、従来のVLA向けオンライン強化学習は、この修正を取り込めないか、違いを区別しない教師情報として扱っていた。実際には、人による修正は一様に雑音を含むわけではなく、動作のある次元では一貫し、別の次元では変動する。 著者らは、凍結したVLAに対する実機の強化学習の枠組みBEEを導入し、人の実演の模倣を超えた方策の改善を目指す。人の修正を再現すべき動作とみなさず、制約に関する証拠として扱う。修正モデルは、あるVLAの提案に人ならどのような修正を加えるか、また各動作次元で修正がどれほど一貫しているかを予測する。この一貫性の予測から、方策の最適化に課す制約の強さを次元ごとに決める。修正が安定している次元では人の動作に近く保ち、ばらつく次元では制約を緩める。 実世界の物体操作3課題とLIBERO-Proのシミュレーション1課題で、オンライン学習に使うデータ量をそろえて評価した。BEEは全課題で最高の成功率を達成し、平均は91.2%で、RLTの57.5%、DSRLの42.1%を上回った。また、実世界の全課題で人の介入率が最も低かった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes human corrections indispensable. However, existing online RL methods for VLAs either cannot incorporate such corrections or fold them into undifferentiated supervision. Yet human corrections are not uniformly noisy but reliable along some action dimensions and variable along others. Building on this, we introduce BEE, an intervention-adaptive framework for real-world RL on a frozen VLA that lets the policy go BEyond Expert imitation. We formulate human corrections not as actions to reproduce but as evidence about a constraint: a Correction Model predicts how a human would correct a given VLA proposal and how consistent the correction is along each action dimension. This predicted consistency sets the per-dimension tightness of a constraint on policy optimization. Where corrections are consistent the policy stays close to the human, and where they vary, the constraint relaxes. We evaluate BEE on three real-world manipulation tasks and one LIBERO-Pro simulation task at a matched online-data budget. BEE attains the highest success rate on every task, 91.2% on average against 57.5% for RLT and 42.1% for DSRL, and the lowest human intervention rate on all real-world tasks.

arXiv ID: 2609.27450 / 要約の誤りについて