arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

人の手の骨格を共通言語に、ロボット操作を学ぶ

Skel-WAM: A Hand-Skeleton-Conditioned World Action Model for Human-to-Robot Manipulation Transfer

Zetao Cai, Yaping Li, Yiqun Wang, Xinyu Zhan, Yuyin Yang, Haoxiang Ma, Kailin Li, Tao Lu, Jiangmiao Pang, Linning Xu, Dahua Lin

この論文をやさしく読む

ひとことで言うと

人とロボットの手を共通の骨格表現に変換し、人の操作動画もロボットの学習に使えるようにする方法です。

何に役立つ?

ロボットだけで多様な実演を集める負担を減らし、人の動画で経験を補う用途が考えられます。評価では、ロボットの学習データにない課題変形で実環境の成功率が改善しました。

この研究の面白いところ

動きを理解・予測する部分を人とロボットで共有し、実際のロボット制御への変換は別の専門モデルに任せます。人の動画に対応するロボット命令を付ける必要がない構成です。

どこまで分かった?

実機の評価は4つの両手作業、シミュレーションは7課題です。79.86%は実環境、63.29%は模擬環境の平均であり、38.89%から86.11%への改善は未学習の課題変形という別の比較です。あらゆる作業への一般化を示すものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ロボットの実演データは収集コストが高く、課題のさまざまな変形を十分にカバーできないことが多い。人の動画は低コストで補完的な操作経験を提供するが、そこから学ぶには、見た目と行動空間における身体の違いを橋渡しする必要がある。本研究では、統一された手骨格の運動インターフェースを通してこの違いをつなぐ世界行動モデル、Skel-WAMを導入する。 中心となる着想は、共通の手のトポロジーを介して人とロボットの動きを整合させることである。そのため、場面内に動きを位置付ける骨格の重ね描きと、手の運動学を明示的に符号化する構造化された2.5次元キーポイントを組み合わせる。Video ExpertとKeypoint ExpertはMixture-of-Transformersを通して視覚と骨格のダイナミクスを共同で学習し、別途ロボットで学習したAction Expertがこれらの予測を実行可能な制御へ変換する。この分離により、人の動画にロボット行動ラベルを付けなくても、人とロボットの実演から共有ダイナミクスを直接教師あり学習できる。 4つの実環境の両手作業と7つのシミュレーション課題で、Skel-WAMはそれぞれ79.86%と63.29%の平均成功率を達成し、最も強いベースラインをそれぞれ22.22パーセントポイント、8.28パーセントポイント上回った。人とロボットの共同学習は、ロボットの学習データにない課題変形について、実環境での成功率を38.89%から86.11%へと2倍以上に高めた。これらの結果は、共有の骨格インターフェースが人とロボットのデータを通じた共同学習を可能にし、補完的な人の実演によってロボットが対応できる課題の範囲を広げることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robot demonstrations are expensive to collect and often provide limited distributional coverage of task variations. Human videos offer a low-cost source of complementary manipulation experience, but learning from them requires bridging embodiment gaps in visual appearance and action spaces. We introduce Skel-WAM, a world action model that bridges these differences through a unified hand-skeleton motion interface. The key insight is to align human and robot motion through a common hand topology, combining skeleton overlays that ground motion in the scene with structured 2.5-D keypoints that encode explicit hand kinematics. Video and Keypoint Experts jointly learn visual and skeletal dynamics through a Mixture-of-Transformers, while a separate robot-trained Action Expert maps these predictions to executable controls. This separation enables human and robot demonstrations to directly supervise shared dynamics without requiring robot action labels for human videos. Across four real-world bimanual tasks and seven simulated tasks, Skel-WAM achieves average success rates of 79.86% and 63.29%, surpassing the strongest baseline by 22.22 and 8.28 percentage points, respectively. Human-robot cotraining more than doubles real-world success on task variations absent from robot training data, from 38.89% to 86.11%. These results demonstrate that a shared skeletal interface enables joint learning across human and robot data and expands robot task coverage through complementary human demonstrations.

arXiv ID: 2609.21514 / 要約の誤りについて