明示的な軌道なしでロボットの動作方策に先読みを学ばせる
Learning Foresight without Explicit Trajectories for 3D Diffusion Policies
この論文をやさしく読む
ひとことで言うと
細かな未来軌道を作らず、操作が向かう方向を小さな潜在表現としてロボット方策に学ばせます。
何に役立つ?
考えられる用途は、既存の3D拡散方策へ少ない追加パラメータで先読み能力を加えることです。
この研究の面白いところ
訓練時だけ未来のグリッパ状態を教師に使い、3.52%のパラメータ増で実機5タスクの成功率を49.0%から72.0%へ改善しました。
どこまで分かった?
報告された複数ベンチマークと実機5タスクでの結果です。推論時に未来の正解状態を与える方式でも、明示的な長期計画を生成する方式でもありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
3次元拡散方策は、現在の観測から幾何学的な根拠を持つ動作を生成することに優れる。しかし、操作を成功させるには、いま可能な運動を知るだけでなく、相互作用がどこへ向かうかを予測する必要がある。従来の方策の多くは、こうした先読みが動作学習から暗黙に生まれることに委ねている。 本研究は、明示的な計画を導入せずに先読みを与える、単純だが効果的な方法Movement Trend Guidanceを導入する。方策は短い観測履歴から、相互作用の進展を表すコンパクトな潜在表現を学ぶ。訓練中は、まばらに抽出した将来のグリッパー状態をこの表現の教師信号とする。推論時には潜在表現だけを残し、現在の観測とともに、未来に関する条件付けに用いる。この潜在表現は動作生成全体の条件となり、追加のゲート付きFiLM分岐はUNetのボトルネックだけで使用する。 DP3に対して追加するパラメータはわずか3.52%だが、元の密な動作列と、先を見て計画を繰り返し更新する移動ホライズンの定式化を保ちながら、RoboTwin2.0、LIBERO-40、DexArtで一貫してDP3を改善する。50課題を混合して訓練するRoboTwin2.0では56.1%に対して62.8%、LIBERO-40では37.08%に対して71.93%、5つの実機ロボット課題では49.0%に対して72.0%に達する。これらの結果は、正確にどこへ動くかを指定しなくても、相互作用の向かう先を知ることが拡散方策に大きな利益をもたらし得ることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
3D diffusion policies are strong at generating geometrically grounded actions from current observations, but successful manipulation requires not only knowing what motion is feasible now, but also anticipating where the interaction is heading. Existing policies largely leave such foresight to emerge implicitly from action learning. We introduce Movement Trend Guidance, a simple but effective way to provide this foresight without introducing an explicit plan. From a short observation history, the policy learns a compact latent representation of interaction evolution. During training, sparse future gripper states supervise this representation; at inference, only the latent is retained as future-oriented conditioning alongside the current observation. The latent provides global conditioning for action generation, while an additional gated FiLM branch is used only at the UNet bottleneck. Despite adding only 3.52% more parameters to DP3, our method preserves the original dense-action and receding-horizon formulation and consistently improves upon DP3 across RoboTwin2.0, LIBERO-40, and DexArt. It reaches 62.8% vs. 56.1% in 50-task RoboTwin2.0 mixed training, 71.93% vs. 37.08% on LIBERO-40, and 72.0% vs. 49.0% on five real-robot tasks. These results show that a diffusion policy can benefit substantially from knowing where an interaction is heading, without being told exactly where to move.
arXiv ID: 2609.20669 / 要約の誤りについて