arXiv論文メモ
新着一覧
cs.CV / cs.AI / cs.RO · 査読状況未確認

潜在世界モデルで失われた物体の動きを回復する学習法

Frozen Flows Forget: Diagnosing and Restoring Lost Motion in a Latent-flow World Model

Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou, Xiaojun Zhu, Houde Liu

この論文をやさしく読む

ひとことで言うと

物体が動かなくなる潜在世界モデルを調べ、復号結果から学習信号を与えて動きの予測を回復した。

何に役立つ?

考えられる用途は、物体操作向け世界モデルの学習と評価の改善である。要旨の結果はモデルの比較評価による。

この研究の面白いところ

問題の原因を潜在表現ではなく、変化の時点を教えない学習信号に特定した点と、画素誤差の評価上の落とし穴が特徴である。

どこまで分かった?

元モデルより予測品質が上がり、大規模条件で参照方法との差をほぼ半分縮めた。要旨には実ロボット作業での性能は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

固定した自己教師ありの潜在空間でフローを積分する潜在世界モデルは、安定して低コストで学習できる。しかし、物体操作にとって最も重要な動きが、気付かないうちに失われる。事前学習したフローでは操作対象の物体がまったく動かず、潜在空間だけの損失でフローを再学習すると、静止が瞬間移動のような動きに置き換わるだけだった。著者らは、この失敗の原因を表現ではなく学習信号に求める。時点の少ないアンカーと潜在空間だけの教師信号では、予測期間のどこで変化すべきかが示されない。 提案するDecode-augmented rollout training(DART)は、表現は固定したまま、復号経路からの教師信号を用いてフローだけを再学習し、この問題を修復する。DARTは全評価手順で潜在空間だけを用いる元のモデルを上回り、動きの時間構造を回復し、予測された動きを場面と再び結び付けた。規模を大きくすると予測品質がさらに改善し、正解情報を使う補間の参照方法との差の残りをほぼ半分埋めた。評価については、画素誤差だけだと静止した予測が有利になるという予想外の結果も報告する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow never moves the manipulated object; retraining it with latent-only losses only trades stillness for teleport-like motion. We trace the failure to the training signal, not the representation: anchor-sparse, latent-only supervision never says where along the horizon change belongs. Decode-augmented rollout training (DART) repairs this while keeping the representation frozen, retraining only the flow with decode-path supervision. DART outperforms its latent only parent on the full protocol, restores the temporal structure of motion, and re-couples predicted motion to the scene; at larger scale it further improves prediction quality, closing nearly half the remaining gap to an oracle-informed interpolation reference. Finally, we report an unexpected finding about evaluation: pixel error alone rewards frozen predictions.

arXiv ID: 2609.28414 / 要約の誤りについて