arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

人型ロボットに回転するカーブキックを学習させる

Banana Kick: Response-Informed Skill Evolution for Humanoid Soccer

Hao E. Zhang, Ruize Geng, Raihan Haque, Khalil Zbiss, Guanyang Luo, Hui-ping Wang, H. Eric Tseng, Ding Zhao

この論文をやさしく読む

ひとことで言うと

人型ロボットが通常のキックから、ボールを回転させて曲げるキックを学ぶ方法。

何に役立つ?

接触を伴う別のロボット技能を、既存の動作から学習させる設計の参考になる。

この研究の面白いところ

報酬を単に強めるのではなく、実際の反応が改善する目標変更だけを採用した。

どこまで分かった?

実機での移行は30回の試行で示された。要旨の数値はこのキック課題と比較設定での結果である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人型ロボットのキックには全身の協調運動と正確な接触が必要で、カーブするキックにはボールの回転と空気力学的な曲がりを生む接触力学が必要である。動作の模倣は通常のキックの確かな出発点となるが、強化学習は基礎となる蹴り方を変えずに速度や狙いの精度だけを高めることがある。この出発点を質的に異なる接触の多い技能へ適応させることは、報酬が密で最適化が安定していても失敗し得る。現行の方策の反応に対して課題の目的が局所的に平坦なときに失敗し、これを一次の学習信号の枯渇と呼ぶ。対処のため、方策適応における閉ループの目的継続法、反応に基づく技能進化(RISE)を提案する。RISEは保存済みの実行から推定した反応感度によって、範囲を限った目標の変更を順位付けし、キックの信頼性を保ちながら、検証された反応の進歩がある場合だけ更新を採用する。解析から、飽和した回転報酬の大きさを変えても回転がゼロのときの一次感度は回復せず、連動した接触反応を適応させれば回転を生む学習可能な道筋ができることが分かった。接触とMagnus力の空気力学を較正し、シミュレーションから実機へ移す人型ロボットのキック手順にRISEを組み込んだ。実験では通常のキックが高回転のカーブキックへ進化し、ボールの平均回転速度11.55ラジアン毎秒、学習進捗に基づくカリキュラムより平均評価得点が19.8%高く、回転と狙いの両方の目標達成率が15.2%から50.9%に上がった。要素除去と反応診断が仕組みを裏付け、モーションキャプチャーで記録した実機試行30回でも学んだカーブキックが一貫して移ったことを示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Humanoid kicking requires coordinated whole-body motion and precise contact, while a banana kick demands contact mechanics that generate ball spin and aerodynamic curvature. Motion imitation provides a reliable ordinary-kick prior, but reinforcement learning may improve shot speed and placement accuracy without changing the underlying kicking technique. Adapting this prior to a qualitatively different contact-rich skill can fail even when the reward is dense and optimization remains stable. The failure occurs when the task objective is locally flat over the current policy's responses. We term this condition first-order learning starvation. To address it, we propose response-informed skill evolution (RISE), a closed-loop objective-continuation method for policy adaptation. RISE ranks bounded objective changes using response sensitivity estimated from cached rollouts and accepts updates only when they produce verified response progress while preserving kicking reliability. Our analysis shows that rescaling a saturated spin reward cannot recover first-order sensitivity at zero spin, whereas adapting coupled contact responses can provide a learnable path to spin generation. We integrate RISE into a humanoid kicking pipeline under calibrated contact and Magnus-force aerodynamics, and sim-to-real transfer. Experiments show that RISE evolves the ordinary kick into a high-spin curved kick with 11.55 rad/s mean ball spin, improves the mean evaluation score by 19.8% over a learning-progress curriculum, and raises joint target attainment from 15.2% to 50.9%. Ablations and response diagnostics support the mechanism, while 30 motion-capture-recorded physical trials demonstrate consistent hardware transfer of the learned curved kick. Project website: https://haozhang-thu.github.io/bananakick/

arXiv ID: 2609.27269 / 要約の誤りについて