arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

人の手の接触を手掛かりにロボットの器用な操作を学ぶ

Dexterous Robot Manipulation from Human Demonstrations via Contact-Anchored Retargeting and Residual Policy Learning

Zihao Yang, Chengyuan Liu, Yu Zhou, Runze Lv, Tianyu Cui, Sheng Yi, Haohua Zhu, Irvine Lu, JieQ Sun

この論文をやさしく読む

ひとことで言うと

人の指の曲がり方をそのまままねるのではなく、物体のどこをどの指で触るかをロボットに引き継ぐ方法です。記録にない接触力はシミュレーションを使って補います。

何に役立つ?

考えられる用途は、人の動作記録から異なる形のロボットハンドを学習させることです。実機データを使わない学習に加え、四つの両手課題の実機実行も報告されています。

この研究の面白いところ

接触構造を移す処理と、物理的に実行できるよう残差方策で直す処理を組み合わせています。手の形態をまたいで使える情報に焦点を当てています。

どこまで分かった?

軌道再構成での成功率と、四つの課題の実機実行は別の結果です。実機試行の成功率や、62.4ポイント改善の比較指標の詳細は要旨に示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

実演から器用な物体操作を学習する際には、データがボトルネックになる。把持が成功するかどうかを決める接触力が、大規模に集められる人の実演データにはいずれも含まれていないためである。本論文は二つの観察に基づく。第一に、人の手からロボットの手へ変えても引き継がれるのは、関節運動そのものではなく、どの指の領域が物体のどの位置にどの順番で触れるかという、実演の接触構造である。第二に、物理的な整合性を課題ごとに設計する必要はない。多様な実演を用いて一度学習した一つの残差強化学習(RL)方策で、運動学的な記録を、物理的に整合し接触情報を付与した軌道へ修正でき、同じ残差の定式化によって、リターゲティング後にも動力学的な実行可能性を回復できる。 これらの観察から、実機ロボットの学習データを使わず、人のモーションキャプチャ記録を器用なロボット方策へ変換する三段階の処理系を構成する。まず、シミュレーション上のMANO手モデルによる物理的精緻化で接触と力を復元する。次に、手の形態に依存しない目的関数を用いた接触基準のリターゲティングで、実演の接触構造を移す。最後に、残差方策学習で結果をロボットの駆動に適応させる。 この処理系は、各設定で一つの共通方策を用いて、25,454本の片手軌道を再構成し、成功率を7.3%から59.3%へ高めるとともに、25種類の両手課題で16.0%から62.4%へ高めた。一つの人のデータセットを形態の異なる四つのロボットハンドへ移し、62.4パーセントポイントの改善を得た。また、実機の学習データを一切使わず、接触を多く伴う四つの両手課題を実機で実行した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Learning dexterous manipulation from demonstrations is bottlenecked by data: the contact forces that determine whether a grasp succeeds are absent from every scalable source of human demonstrations. This paper builds on two observations. First, what survives the change from a human hand to a robot hand is the contact structure of a demonstration - which finger regions touch which object locations, and in what order - rather than its joint motion. Second, physical consistency need not be engineered per task: a single residual reinforcement learning (RL) policy, trained once across diverse demonstrations, can repair kinematic recordings into physically consistent, contact-annotated trajectories, and the same residual formulation restores dynamic feasibility after retargeting. These observations yield a three-stage pipeline that converts human motion-capture recordings into dexterous robot policies with no real-robot training data: physics refinement with a simulated MANO hand recovers contacts and forces, contact-anchored retargeting transfers the demonstrated contact structure through an objective independent of hand morphology, and residual policy learning adapts the result to robot actuation. The pipeline reconstructs 25,454 single-hand trajectories (success 7.3% -> 59.3%) and 25 dual-hand tasks (16.0% -> 62.4%) with one shared policy per setting, transfers one human dataset to four morphologically distinct robot hands (+62.4 pp), and executes four contact-rich bimanual tasks on physical hardware with zero real-robot training data.

著者のコメント

16 pages, 6 figures, 5 tables. Technical report. Code: https://github.com/DexGEM-Lab/real2sim2real

arXiv ID: 2609.24093 / 要約の誤りについて