人の介入で模倣学習ロボットの修正動作を学ぶ
Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
この論文をやさしく読む
ひとことで言うと
模倣学習したロボットの動作に、人の修正を使って小さな補正を上乗せする学習法を示した。
何に役立つ?
実演を大量に集めず、接触を伴う操作をオンラインで改善する方法として検討できる。
この研究の面白いところ
人の介入を、現在の修正動作の教師信号と、それ以前の動作への報酬の両方に使う。
どこまで分かった?
性能比較は要旨に記載された五つの操作課題と10分のオンライン学習条件での結果。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
模倣学習では実演からロボットが操作を覚えられるが、得られた方策は学習データから外れると失敗しうる。実演を追加で集めるには人手が大きくかかる。人が関与する強化学習はオンライン学習中の修正情報を利用するものの、事前学習した模倣方策を改良するのではなく、作業全体の方策を学ぶことが多い。本研究は、固定した模倣方策の上に修正動作を重ねて学ぶ、人が関与する残差強化学習の枠組みRes-HILを提案する。人が介入するたびに、残差方策への直接の教師信号と、その前の自律動作に対する報酬の形づけという二種類の学習信号を得る。これらを残差方策のゼロ初期化と組み合わせ、オンライン学習を安定させ、速める。 高精度な操作と長い手順の操作を含む、接触の多い五つの課題で評価した。最初の実演がわずか20件でも、10分間のオンライン学習後には、比較した最先端の全方策を学ぶ人間参加型強化学習法と、人の誘導を使わない残差微調整法をすべての課題で上回った。また事前学習済みの基礎方策を改善し、5倍の実演で学んだ模倣方策も上回った。構成要素の検証では、残差への直接の教師信号が性能に不可欠で、人の介入を考慮した報酬の形づけが学習効率を大きく改善した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.
arXiv ID: 2609.30023 / 要約の誤りについて