人の手の動作をロボットの手へ効率よく移す学習法
FlashDexRetarget: Accelerating Dexterous Manipulation Data Generation through Multi-Motion Retargeting
この論文をやさしく読む
ひとことで言うと
人が物を扱う動きを、形の異なるロボットの手で実行できる動作へ変換する強化学習の方法です。
何に役立つ?
考えられる用途は、人の操作記録を再利用してロボットの器用な操作用データを増やすことです。評価では、比較対象より高い変換成功率と少ない学習計算量を報告しています。
この研究の面白いところ
手の姿勢だけでなく、物体の形、手との距離、将来の軌道を学習に与えます。左右の手を別ネットワークで扱い、1,000動作まで規模を広げた点も特徴です。
どこまで分かった?
90%の成功率は50動作のベンチマークでの値です。計算量の最大100倍の差も、評価された比較手法との関係です。現実の実演を用いた定性的再生は報告されていますが、要旨は実機での広範な操作成功率までは示していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
人の手と物体の操作実演は、器用なロボット操作のために再利用できるデータ源となるが、異なる身体構造へ転用するには、物理的に実行可能な動作への変換が必要である。既存の物理ベース手法には、動作変換の成功率、動作ごとの学習効率、またはその両方に限界がある。これらに対処するため、高い成功率と効率を備える器用な動作変換の強化学習フレームワーク、FlashDexRetargetを提案する。実演された相互作用を学習しやすくするため、物体の点群観測、手と物体の距離特徴、将来軌道の符号化を組み合わせ、物体の動きと参照となる手・物体間の関係を指導する相補的な報酬を用いる。さらに学習を高速化するため、左右の手に個別のアクター・クリティックネットワークを用い、オフポリシーアルゴリズムFlashSACを器用な動作追跡に適応させる。 単一物体および2物体との相互作用を含む50動作のベンチマークでは、FlashDexRetargetは90%の成功率を達成した。これは評価したサンプリングベースの比較手法の約2.5倍であり、学習に必要な計算量は、評価した強化学習ベースの比較手法に比べて最大で100分の1となった。XHandとSharpa Wave Handの両方で一貫した改善が見られ、構成要素別のアブレーションによって各設計の寄与を調べた。50動作のベンチマークに加え、200、500、1,000動作の実験から、規模を拡大しても手法が安定しており、学習集合が増えるにつれて変換に成功する動作をより効率的に生成できることが示された。現実世界で収録した実演を使った定性的な再生結果も、記録された人の操作に本フレームワークを適用できることを例示している。動画とコードは https://davian-robotics.github.io/FlashDexRetarget/ で公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Human hand-object demonstrations offer a reusable source of dexterous robot manipulation data, but transferring them across embodiments requires physically feasible retargeting. Existing physics-based approaches face limitations in retargeting success, motion-specific training efficiency, or both. To address these limitations, we introduce FlashDexRetarget, an RL-based framework for high-success, efficient dexterous motion retargeting. To make the demonstrated interaction easier to learn, we combine object point-cloud observations, hand-object distance features, and future trajectory encodings with complementary rewards that supervise object motion and reference hand-object relationships. To further accelerate learning, we employ separate left- and right-hand actor critic networks and adapt the off-policy algorithm, FlashSAC to dexterous motion tracking. On a benchmark of 50 motions spanning single-object and two-object interactions, FlashDexRetarget achieves a 90% success rate, approximately 2.5x that of the evaluated sampling-based baselines, while requiring up to 100x less training compute than the evaluated RL-based baselines. Evaluations on both XHand and Sharpa Wave Hand show consistent gains, and component-wise ablations examine the contributions of our design choices. Beyond the 50-motion benchmark, experiments with 200, 500, and 1,000 motions demonstrate that our method remains stable at larger scales and produces successful retargeted motions more efficiently as the training set grows. Qualitative replay results using real-world-captured demonstrations further illustrate the applicability of our framework to recorded human manipulation. Videos and code are available at https://davian-robotics.github.io/FlashDexRetarget/
arXiv ID: 2610.01849 / 要約の誤りについて