他の身体の動きから学ぶ世界モデルでロボット方策を補助
Latent Policy Steering: An Efficient and Flexible Framework for Cross-Embodiment Transfer
この論文をやさしく読む
ひとことで言うと
別のロボットや人の動きから世界の変化を学び、対象ロボットの少数の実演を使って既存の行動方策を補助します。
何に役立つ?
新しい身体のロボットで実演収集の負担を抑える用途が考えられます。既存方策を再学習せずに補助できる設計です。
この研究の面白いところ
行動空間を直接そろえるのではなく、画像で見える世界の動きに注目します。テスト時にも世界モデル内で探索して方策を誘導します。
どこまで分かった?
50件は対象身体の実演数で、事前学習全体のデータ量ではありません。改善値は相対的な性能向上で、実世界の成功率が62ポイント増えたという意味ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
学習したロボットの視覚運動方策の性能は学習データの量と質に強く依存するが、実世界のロボットで高品質な実演を集めるには依然として費用がかかる。大規模なロボットや人間のデータセットは増えているものの、身体構造の隔たりと行動空間の不一致により、そのまま活用することは難しい。他の身体の経験を再利用して対象の身体の学習を改善する身体間転移は、ロボットごとのデータ収集を超えて学習を拡大するために重要である。 本研究では、身体をまたいで共通する、運動に対して世界がどう応答するかという視覚的ダイナミクスから学び、少量の対象身体データをテスト時に有効利用することで、効率的な転移が可能になることを見いだす。提案するLatent Policy Steering(LPS)は、身体に依存しない事前学習段階を設け、多様な身体のオプティカルフローを使って画像ベースの世界モデル(WM)を学習する。得られたWMを、ロボットの行動を用いて対象の身体に微調整する。続いて、微調整データから大きく離れない計画をWMの潜在空間で探索し、基礎方策をよりよい行動へ誘導する。 LPSは方策に依存しない枠組みであり、方策自体を再学習せずに異なる方策へ柔軟に対応できる。未見の対象身体でわずか50件の実演を用い、Robomimicと実世界の評価において、Diffusion Policyの平均性能をそれぞれ相対的に16%と62%、Pi0.5を8%と14%向上させた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The performance of learned robot visuomotor policies depends heavily on the size and quality of their training data, yet collecting high-quality demonstrations remains costly for robots in the real world. Although large-scale robot and human datasets are increasingly available, embodiment gaps and mismatched action spaces make them difficult to leverage directly. Cross-embodiment transfer, reusing experience from other embodiments to improve learning on a target embodiment, is therefore crucial for scaling robot learning beyond per-robot data collection. In this work, we find that efficient transfer can be achieved by learning from what is shared across embodiments, the visual dynamics of how the world responds to motion, and by effectively exploiting the scarce target-embodiment data at test time. The proposed framework, called Latent Policy Steering (LPS), implements an embodiment-agnostic pretraining phase, which trains an image-based World Model (WM) with optical flow across diverse embodiments. The resulting WM is finetuned on the target embodiment with robot actions. It then steers the base policy toward better actions by searching in the WM's latent space for plans that stay close to the finetuning data. LPS is a policy-agnostic framework: it can flexibly accommodate different policies without having to retrain them. In Robomimic and real-world evaluations, LPS improves the average performance of Diffusion Policy relatively by 16% and 62%, and Pi0.5 by 8% and 14%, with only 50 demonstrations on an unseen target embodiment.
arXiv ID: 2609.22521 / 要約の誤りについて