映像の変化を潜在表現に残す世界モデルMotionJEPA
MotionJEPA: Preventing Temporal Feature Collapse by Capturing Visual Changes in Latent Space
この論文をやさしく読む
ひとことで言うと
動画から学ぶモデルが背景など変化の少ない情報に偏らず、動きの情報も内部表現に残せるようにする学習方法です。
何に役立つ?
静止した背景に注意を奪われやすい環境で、動きを踏まえた計画に使う世界モデルの表現を改善する用途があります。
この研究の面白いところ
画像を画素単位で復元するのではなく、時間差分画像の埋め込みを予測します。動的項は表現の分布を指定せず、動きの情報が存在することだけを促します。
どこまで分かった?
計画性能の改善は静的背景の妨害要素を含む4環境で示されています。要旨には具体的な成功率や、実機での評価結果は記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Joint Embedding Predictive Architecture(JEPA)は、画像を再構成せずに、特定の課題に依存しない潜在世界モデルを学ぶ有望な枠組みである。しかし標準的なJEPAの学習には、変化の遅い特徴を優先する強い帰納バイアスがあり、特徴の抑制と潜在表現の崩壊を引き起こす。逆ダイナミクスは時間方向の崩壊を防ぐが、行動ラベルに依存し、ラベルのない一般的なダイナミクスを埋め込む動機付けは弱い。 本研究では、Difference Image and Single image embedding Regularization(DISReg)という新しい正則化法を導入する。これは、画素の再構成損失を一切使わずに時間差分画像の埋め込みを予測する、逆ダイナミクス型モジュールに基づき、静的・動的特徴をバランスよく学ぶよう促す。 DISRegは静的項と動的項からなる。静的項は画像埋め込みの分布を整え、変化の遅い特徴を促す。一方の動的項は、埋め込みへの直接的な正則化とは異なり、画像埋め込みの形や分布には制約を課さず、動的な特徴が含まれることだけを促す。この正則化法を標準的なJEPAに組み込み、新たなアーキテクチャMotionJEPAを構成する。 潜在表現のプロービングにより、MotionJEPAは他の手法よりも情報を幅広く含む表現を生成することが示された。軌道解析では、曲率の低い、幾何学的に単純な潜在埋め込みを保つことが分かった。さらに4環境にわたり、静的背景の妨害要素がある条件で、下流の計画課題の成功を改善することを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Joint Embedding Predictive Architectures (JEPAs) are a promising paradigm for learning task-agnostic latent world models without visual reconstruction. However, standard JEPA training exhibits a strong inductive bias towards slow features, causing feature suppression and the collapse of latent representation. While inverse dynamics provides temporal anti-collapse, it relies on action labels and offers little incentive to embed general, unlabeled dynamics. We introduce Difference Image and Single image embedding Regularization (DISReg), a novel regularizer that builds on an inverse-dynamics-style module that predicts temporal difference image embeddings without any pixel reconstruction loss, encouraging balanced static and dynamic feature learning. DISReg consists of a static term that shapes the distribution of the image embedding and encourages slow features, and a dynamic term, which, unlike direct regularization on the embedding, imposes no constraint on the image embedding's shape or distribution and instead only incentivizes that dynamic features be present. By integrating this regularizer into a standard JEPA, we establish our new architecture, MotionJEPA. Latent probing demonstrates that MotionJEPA produces more complete representations than other methods, and our trajectory analysis shows it maintains geometrically simple latent embeddings with low curvature. We further show that MotionJEPA improves downstream planning success under static-background distractors across four environments.
著者のコメント
*Equal contribution (Markus Karmann, Shile Li). Code available: https://github.com/mkarmann/motion-jepa
arXiv ID: 2609.23881 / 要約の誤りについて