接触の結果を重視してロボットの捕球動作を選ぶ
Outcome-Sensitive Motion Search for Impact-Aware Dexterous Catching
この論文をやさしく読む
ひとことで言うと
接触直前の小さな動きで結果が大きく変わることに着目し、ロボットの捕球実演を選び直します。
何に役立つ?
衝撃を抑えたロボットの捕球方策を学習させる際に役立つ可能性があります。
この研究の面白いところ
教師の失敗例も動作探索で補い、生徒の行動誤差を考慮した完全な軌道で実演候補を検証します。
どこまで分かった?
成果は要旨ではシミュレーション実験で示されています。実機での捕球結果は記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
熟練した人は、動く物体との接触位置、速度合わせ、接触後の追従動作を協調させ、速い物体も衝撃を抑えて捕らえられる。しかし、強化学習で衝撃を考慮した捕球を学ぶのは難しい。方策には確実な接触と把持に加え、接触へ移る敏感な段階を調節する能力が必要だからである。また、多くの状態情報を利用できる強化学習の教師が有能でも、実機で使う模倣学習の生徒にとって理想的な実演を必ずしも作れない。教師の失敗は課題の対象範囲を狭め、接触前の動きの小さな違いが衝撃や把持の結果を大きく変える。 本研究はこの現象を介入による結果感度として特徴付け、実演を重点的に作るための結果感度の高い時間窓(OSW)を導入する。この定式化に基づき、課題条件ごとに成功するOSWの動きの多様体を学習し、局所的な測地線探索で教師の成功軌道を改良するとともに、教師が失敗する課題条件を補う「Outcome-Sensitive Motion Search」を提案する。次に、模倣学習の生徒の行動誤差を較正したモデルを用い、候補動作を完全な実行軌道で検証し、成功した実行だけを実演として残す。広範なシミュレーション実験では、教師が失敗する条件を効果的に補い、得られた模倣学習方策が捕球成功率と衝撃の軽減の両方で、多くの状態情報を利用する強化学習の教師を上回った。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Skilled humans can catch fast-moving objects softly by coordinating interception, velocity matching, and follow-through to mitigate impact. Learning such impact-aware catching with reinforcement learning (RL), however, is challenging, as the policy must achieve reliable interception and grasping while regulating the sensitive transition into contact. Moreover, even a capable privileged-state RL teacher may not provide ideal demonstrations for a deployable imitation-learning (IL) student: teacher failures limit task coverage, while small variations in pre-contact motion can produce substantially different impact and grasping outcomes. We characterize this phenomenon through interventional outcome sensitivity and introduce the outcome-sensitive window (OSW) to guide targeted demonstration construction. Building on this formulation, we propose Outcome-Sensitive Motion Search, which learns a task-conditioned manifold of successful OSW motions and performs local geodesic search to refine successful teacher rollouts and repair task conditions where the teacher fails. We then validate candidate motions through complete rollouts under a calibrated IL-student action-error model and retain only successful executions as demonstrations. Extensive simulation experiments demonstrate that our method effectively repairs task conditions where the teacher fails and enables the resulting IL policy to outperform the privileged RL teacher in both catching success and impact mitigation.
arXiv ID: 2609.29020 / 要約の誤りについて