arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

骨格表現で異なる形のロボットへ操作経験を移す

SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation

Pengjun Niu, Yujia Xie, Rui Peng, Hang Zhao, Ke Liu

この論文をやさしく読む

ひとことで言うと

関節の構成が違うロボットでも共有できる骨格の表し方を使い、一方で学んだ動作をもう一方へ移します。

何に役立つ?

ロボットが変わるたびに課題の実演を集め直す負担を減らす用途が考えられます。移転先の実演や方策更新を使わない設定で評価されています。

この研究の面白いところ

共通の25次元表現を知覚と行動の両方で使い、最後だけ各ロボット用の制約付き変換を行います。連続体ロボットへの実機適用も報告しています。

どこまで分かった?

43.3%はLIBERO-Cross10の1,000エピソードでの成功率です。実機の三課題には別の成功率が要旨に示されていないため、この数値を実機の成績と読むことはできません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

異なる身体構造のロボット間で操作経験を再利用することは、ロボット学習を拡大し、課題ごとのデータ収集の繰り返しを減らすために重要である。しかし身体構造が変わると、見た目、行動の次元と意味、同じツール姿勢を実現する全身の構成が変わる。本研究では、単一の学習元から異なる身体構造へ操作を移すため、知覚と制御を一つの明示的な幾何表現で結ぶ、骨格に基づく世界・行動モデルSkelWAMを提案する。 アーム中心線の幾何、ツール中心点(TCP)の姿勢、平行二爪グリッパーの指令で、共通の25次元状態を作る。同じ定義を、標準化した第三者視点・手首視点の観測と、将来の全身動作の目標に用いる。予測的な視覚教師信号で学習した動画・行動の混合Transformerが、標準化した骨格動作のまとまりを予測し、身体構造ごとの制約付きデコーダがそれを関節制御または連続体ロボットの制御に変換する。関節の一対一対応は不要で、移転先課題の実演も、移転先での方策更新も使わない。 学習元のデータだけを用いる身体間転移ベンチマークLIBERO-Cross10を導入する。これは四つの形態群にわたる10種類の移転先身体構造と10課題を含む。このベンチマークで、Frankaで学習したSkelWAMは1,000エピソードで成功率43.3%を達成し、評価した最良のベースラインを36.2ポイント上回った。さらに、JAKA mini2で学習した方策をFeagine A03連続体ロボットでの三つの卓上操作課題に適用し、実環境で身体構造を越えて操作を移す可能性を示す。プロジェクトページ:http://www.liukepku.com/skelwam/index.html

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Reusing manipulation experience across robot embodiments is important for scaling robot learning and reducing repeated task-specific data collection. However, changes in embodiment alter visual appearance, action dimensionality and semantics, and the whole-body configurations that can realize the same tool pose. We present SkelWAM, a skeleton-guided world-action model that couples perception and control through one explicit geometric representation for single-source cross-embodiment manipulation. Arm centerline geometry, tool-center-point (TCP) pose, and parallel-jaw commands form a shared 25-D state. The same definition underlies canonical third-person and wrist observations and future whole-body action targets. Trained with predictive visual supervision, a video-action mixture of transformers predicts canonical skeleton action chunks, which embodiment-specific constrained decoders convert into joint or continuum-robot controls. This formulation requires no one-to-one joint correspondence and uses no target-task demonstrations or target policy updates. We introduce LIBERO-Cross10, a source-only cross-embodiment transfer benchmark covering ten tasks and ten target embodiments across four morphological groups. On this benchmark, Franka-trained SkelWAM achieves 43.3% success over 1,000 episodes, exceeding the best-performing evaluated baseline by 36.2 percentage points. We further deploy a JAKA mini2-trained policy on the Feagine A03 continuum robot for three tabletop manipulation tasks, illustrating the approach's potential for real-world cross-embodiment manipulation. Project page: http://www.liukepku.com/skelwam/index.html

arXiv ID: 2609.21983 / 要約の誤りについて