JEPAの表現学習では潜在空間の形に合う分布が重要
Beyond Gaussian Worlds: Latent Geometry Matters for JEPAs
この論文をやさしく読む
ひとことで言うと
AIがデータの内部構造を学ぶとき、正規分布へそろえればよいとは限らず、元の構造が球面や輪の形なら、それに合う分布が有利だという研究です。
何に役立つ?
JEPAで表現の崩壊を防ぐための目標分布を選ぶ際の理論的な手掛かりになります。潜在構造を線形に読み出しやすく保つことに関係します。
この研究の面白いところ
従来の正規分布の一意性を否定するのではなく、その結論がEuclid的仮定に依存することを示しています。球面上では別の復元保証と、より鋭い近似誤差の上界を得ています。
どこまで分かった?
理論保証には潜在幾何や正例ペアの動力学、厳密な分布整合の条件があります。実験での優位性も最適化が成功する場合の結果です。現実の任意のデータで球面分布が最良という主張ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
近年のJoint-Embedding Predictive Architecture(JEPA)は、学習した表現が等方的Gaussian分布や超球面上の一様分布など、あらかじめ定めた目標分布に従うよう制約することで、表現の崩壊を防ぐ。Klindtら(2026)は、彼らのEuclid的な仮定の下では、Gaussian目標への分布整合によってGaussian潜在変数を線形変換の違いを除いて復元でき、Gaussian分布がこの保証を持つ唯一の分布であることを示した。 本研究では、その解析を、埋め込まれたRiemann多様体上に分布する潜在変数へ拡張する。整列と厳密な分布整合によって線形復元が保証されるような、潜在空間の幾何と正例ペアの動力学に関する条件を導く。特に、潜在変数が球面上に一様に分布し、表現を同じ球面分布に整合させる場合、すべての最適表現は直交変換の違いを除いて潜在状態を復元する。これは、Gaussian分布の一意性が、分布整合型JEPAの普遍的な性質ではないことを示す。非Euclid的な潜在幾何では、線形復元可能な別の分布も認められる。 さらに、Gaussianの世界より球面の世界で厳密に鋭くなる近似復元の誤差上界を導く。Gaussian、球面、トーラスの潜在空間での実験では、最適化が成功する場合、幾何と適合した目標分布はより良い線形復元をもたらす一方、不適合な目標分布は潜在構造を歪めることを示す。この優位性は、高次元のCliffordトーラスの世界でも維持される。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent Joint-Embedding Predictive Architectures (JEPAs) prevent representation collapse by constraining learned representations to follow a prescribed target distribution, such as an isotropic Gaussian or the uniform distribution on a hypersphere. Klindt et al. (2026) showed that, under their Euclidean assumptions, matching a Gaussian target can recover Gaussian latent variables up to a linear transformation, and that the Gaussian is the unique distribution with this guarantee. We extend their analysis to latent variables supported on embedded Riemannian manifolds and derive conditions on the latent geometry and positive-pair dynamics under which alignment and exact distribution matching guarantee linear recovery. In particular, when the latent variables are uniformly distributed on a sphere and the representations are matched to the same spherical distribution, every optimal representation recovers the latent state up to an orthogonal transformation. This shows that Gaussian uniqueness is not a universal property of distribution-matched JEPAs: non-Euclidean latent geometries can admit other linearly recoverable distributions. We further derive an approximate-recovery bound that is strictly tighter for the spherical world than for the Gaussian world. Experiments on Gaussian, spherical, and toroidal latent spaces show that geometrically compatible targets yield better linear recovery when optimization succeeds, whereas mismatched targets distort the latent structure. This advantage persists in high-dimensional Clifford-torus worlds.
arXiv ID: 2609.21656 / 要約の誤りについて