arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

動作の類似性を学び、異なるロボットへ技能を移す

Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision

Maxime Alvarez, Renzo Caballero, Tatsuya Matsushima, Yusuke Iwasawa, Yutaka Matsuo

この論文をやさしく読む

ひとことで言うと

違うロボットが行う似た動作を、共通の潜在表現でも似たものとして学ぶ方法です。

何に役立つ?

一方のロボットの実演を他の身体構成のロボットへ転用する際、機体固有の関節操作へ表現が寄りすぎる問題を減らす用途があります。

この研究の面白いところ

動作を直接予測させず、二つの動作系列の類似度を潜在動作の類似度へ合わせます。関節空間より手先の動きで類似を測り、二機体間も比較する方式が最良でした。

どこまで分かった?

RoboTwin 2.0の二種類の双腕ロボットによる制御された評価です。潜在動作の予測で転移成功率が二倍超となりましたが、要旨には実機での一般化結果はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

汎用ロボット方策がウェブ規模の事前学習から視覚と言語の能力を獲得する一方、実演データの収集には依然として費用がかかり、データは記録したロボットに結び付いている。潜在行動モデル(LAM)は、行動ラベルのない動画から、異なる身体構造のロボット間で共有できる潜在行動を学習することで、この両方に対処する。しかし実際には、LAMは背景の視覚的ノイズに敏感で、異なる2台のロボットの同じ動きを異なる潜在表現に符号化する場合がある。背景ノイズへの対処法の一つは、潜在行動からロボットの行動を予測する補助損失を加えることだが、これは潜在行動空間を身体構造固有の行動空間にさらに結び付ける。 本研究では、同じラベルを動作の類似性による教師信号として用いる別の方法を検討する。任意の2つの潜在行動の類似度が、対応する正解のロボット行動系列同士の類似度と一致するように学習する。LAM自身は正解行動を一切予測しないため、潜在行動に身体構造固有の情報を符号化する必要がない。 RoboTwin 2.0を用い、条件を統制して身体構造間の転移を評価した。2台の双腕ロボットが互いに重ならないタスク群を実演し、すべての実演で一つの方策を学習させた後、それぞれのロボットを、相手だけが実演したタスクで閉ループ評価する。方策の構造とハイパーパラメータ、データセット、評価手順を固定した条件で、正解行動の代わりに潜在行動を予測すると、身体構造間の転移の成功率は2倍を超えた。同じ正解行動を利用する場合でも、類似性による教師信号は、LAMの学習中に正解行動を予測する補助損失より良好に転移する。関節空間の運動ではなく手先の運動から類似度を計算し、2台のロボットをまたいで潜在行動を比較する損失を用いた方法が、本研究で最も良い結果を得た。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAMs are sensitive to background visual noise, and the same motion from two different robots may be encoded with different latents. One solution to the background visual noise is to add an auxiliary loss predicting the robot action from the latent action, further associating the latent action space to the embodiment specific robot action space. We study a different use of the same labels, through action-similarity supervision. The similarity between any two latent actions is trained to match the similarity of the two ground-truth robot action sequences. The ground-truth actions are never predicted by the LAM, so the latent action does not need to encode embodiment specifics. We evaluate cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated. With the policy architecture and its hyperparameters, the dataset, and the evaluation protocol fixed, predicting latent actions instead of ground-truth actions more than doubles cross-embodiment success. Given the same ground-truth actions, similarity supervision transfers better than an auxiliary loss that predicts the ground-truth action during the LAM training. Computing the similarities on end-effector motion rather than joint-space motion, and letting the loss compare latent actions across the two robots, gives the best approach of the study.

arXiv ID: 2609.19846 / 要約の誤りについて