arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

状態間の往復時間を保つ表現で行動計画を学ぶ

Learning Commute-Time-Preserving World Models for Planning

Michael Hauri, Peter Buttaroni, Fabian A. Mikulasch and Friedemann Zenke

この論文をやさしく読む

ひとことで言うと

状態同士の内部的な距離が、環境を行き来する時間と対応するように世界モデルを学ぶ研究です。

何に役立つ?

目標に近づく行動を潜在空間で選ぶ際、単なる見た目の似かよりも移動の難しさを反映した距離を使う助けになります。

この研究の面白いところ

表現を均等に広げる通常の崩壊防止が、計画に必要な距離の尺度を壊す場合があると指摘し、別の正則化で対処します。

どこまで分かった?

回復の証明は可逆・決定論的な動力学と予測器の固定点という条件付きです。性能比較は数値シミュレーションであり、実機での検証は要旨にありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

世界モデルにより、エージェントは、与えられた目標状態までの距離を最も縮める行動列を選び、潜在空間で計画を立てられる。このため、潜在表現の距離が環境内の往復時間を反映していれば、計画に役立つ。グラフラプラシアンのスペクトル埋め込み空間は、固有値に依存する特定のスケーリングを満たす場合に、そのような表現を与える。しかし、大規模で連続的な環境でグラフラプラシアンを具体的に構成することは計算上困難である。自己教師あり学習は、そのような往復時間を保つ埋め込みを大規模に得る自然な方法となる。 ただし本研究では、表現の崩壊を防ぐため一般に等方的な表現を促す既存手法が、「正しい」固有値依存のスケーリングを損ないやすく、その結果、往復時間の表現が不正確になることを示す。この問題に対し、潜在空間の変位予測器と、崩壊を防ぐ対数行列式の正則化を組み合わせた、往復時間保存世界モデル(CTWM)を導入する。可逆で決定論的な動力学の下、かつ予測器の固定点において、正しくスケーリングされたラプラシアン表現を回復することを証明する。数値シミュレーションでは、複雑で連続的な複数の目標到達ベンチマークにおいて、CTWMはタスクに依存しないベースラインLeWMと同等以上の性能を示しながら、パラメータ数を半分に抑える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

World models allow agents to plan in latent space by choosing a sequence of actions that most reduces the distance to a given goal state. Thus, planning can benefit from latent representations whose distances mirror commute-times in the environment. The spectral embedding space of the graph Laplacian provides such a representation, if it obeys a specific eigenvalue-dependent scaling. Unfortunately, instantiating the graph Laplacian is intractable in large, continuous environments. Self-supervised learning offers a natural route to such commute-time-preserving embeddings at scale. However, here we show that existing methods, which commonly encourage isotropic representations to prevent representational collapse, tend to degrade the "correct" eigenvalue-dependent scaling, leading to an inaccurate representation of commute times. To address this problem, we introduce Commute-Time-Preserving World Models (CTWMs), combining a latent displacement predictor and a log-determinant regularizer that prevents collapse, which provably recover the correctly scaled Laplacian representation under reversible deterministic dynamics and at the predictor's fixed point. In numerical simulations, CTWM matches or outperforms LeWM, a task-agnostic baseline, on several complex, continuous goal-reaching benchmarks, while using half the parameters.

arXiv ID: 2610.01373 / 要約の誤りについて