arXiv論文メモ
新着一覧
cs.RO / cs.LG · 査読状況未確認

行動の実行結果を学んで世界モデルの選択精度を改善

D-JEPA: A Decision-Aligned Latent World Model

Shuaijun Liu, Chengyu Wu, Qifu Wen, Feiyang You, Chenglong Zhang, Shuyang Hao, Xi Lin, Ningxin Su

この論文をやさしく読む

ひとことで言うと

予測した未来が目標に近く見えるかだけでなく、実際にどの行動が良い結果を生んだかを学び、候補の選び方を改善します。

何に役立つ?

考えられる用途は、世界モデルを使うロボット制御や行動計画の選択精度の改善です。要旨ではベンチマークと実ロボットの両方で評価しています。

この研究の面白いところ

予測表現を最初から作り直すのではなく、候補間の順位に関する情報で調整し、潜在距離による既存の計画方式でも利用できる形にします。

どこまで分かった?

PushTの87.89%は成功率ですが、RoboTwinと実機の改善はそれぞれ15.04ポイント、17ポイントという表記です。相対的な改善率に読み替えることはできません。要旨には全評価条件や比較対象の詳細はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

潜在世界モデルは行動の結果を予測するが、正確な予測ができても、潜在空間の距離がどの候補の実行成功につながるかを反映するとは限らない。本研究では、意思決定の局所における予測の隔たりを特定する。すなわち、実行候補として競合する少数の未来の中で、目標に近いと予測された候補が、利用可能な別の候補より悪い実際の結果を生む場合がある。 そこで、実行した結果から候補となる未来同士の意思決定に関係する関係性を学ぶ、意思決定に整合した潜在世界モデルD-JEPAを導入する。有界で置換同変な演算子が、目標に対する予測特徴と順序情報を併せて推論し、行動選択の影響が特に大きい箇所で、事前学習された予測の幾何構造を調整する。制限した予測器の適応と共通の順序インターフェースにより、相補的な予測の幾何構造の間にも、この整合を拡張する。さらに、学習した意思決定構造をJEPA互換の未来表現として実現し、標準の潜在距離に基づく計画を通じて導入できるようにする。 潜在空間での制御、物体操作、事前学習済みの行動生成モデル、実ロボット、自動運転にわたる評価で、行動選択の改善を示した。具体的にはPushTで87.89%の成功率、RoboTwinで平均15.04ポイントの改善、実ロボット課題で17ポイントの改善を得た。これらの結果は、意思決定に関係する関係構造が、予測的な世界モデルと有効な制御を直接結び付けることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Latent world models predict the consequences of actions, but accurate prediction does not guarantee that latent distance reflects which candidate will execute successfully. We identify a decision-local prediction gap: among the few futures competing for execution, a candidate predicted closer to the goal can produce a worse realized outcome than an available alternative. We introduce D-JEPA, a decision-aligned latent world model that learns decision-relevant relations among candidate futures from executed outcomes. A bounded, permutation-equivariant operator jointly reasons over goal-relative predictive features and ordinal evidence, refining pretrained predictive geometry where action choices are most consequential. Restricted predictor adaptation and a shared ordinal interface extend this alignment across complementary predictive geometries. D-JEPA further realizes the learned decision structure in JEPA-compatible future representations, enabling deployment through native latent-distance planning. Evaluations across latent control, manipulation, pretrained action-producing models, physical robots and autonomous driving demonstrate improved action selection, including 87.89% success on PushT, a 15.04-point average gain on RoboTwin, and a 17-point gain on physical robot tasks. These results establish decision-relevant relational structure as a direct bridge between predictive world modeling and effective control.

著者のコメント

26 pages, including references and appendices. Project website: https://nebulis-lab.com/D-JEPA

arXiv ID: 2609.24749 / 要約の誤りについて