arXiv論文メモ
新着一覧
cs.LG / cs.CV / math.AT · 査読状況未確認

AIモデルの表現は距離より関係構造で近づく

What Converges in the Platonic Representation Hypothesis? Structure over Geometry

Junwon You, Mihyun Jang, Sangwoo Mo, Jae-Hun Jung

この論文をやさしく読む

ひとことで言うと

異なるAIモデルの表現が似てくるかを調べ、標本間の関係は似ても距離の値までは一致しにくいと示した。

何に役立つ?

モデル表現の類似性を評価するとき、局所か全体かと、関係か距離かを分けて考える助けになる。

この研究の面白いところ

局所・全体と関係・距離を2×2で比較し、画像と言語に加えて動画と文章でも同じ傾向を見た。

どこまで分かった?

結果は要旨で調べた視覚言語モデルと動画文章表現に基づき、あらゆるモデルでの収束を示すものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

プラトン的表現仮説は、能力が高まるモデルの表現が共通のものへ収束することを示唆する。最近の研究はこの主張を、局所的な近傍関係の共有に狭め、いくつかの全体的な類似度指標で見られた能力に応じた傾向は、較正後にはほぼ消えると報告した。本研究は、以前の局所・全体の比較では、局所か全体かという構造の尺度と、何を比較するかが混同されていると指摘する。後者は、どの標本が関係するかという関係構造と、距離・類似度・相関など量的な関係を表す計量幾何とに分かれる。この二要因を分けるため、局所と全体の両方で関係構造と計量幾何を評価する、制御された2×2の枠組みを作る。相互k近傍に対応する全体的な指標としてH₀骨格の重なりを導入し、距離を考慮した対応指標も用いる。画像と言語のモデルでは、較正後の関係構造は両尺度で頑健な表現の収束を示す一方、距離の一致を厳しく求めるほど対応はかなり弱まり、能力依存の傾向も次第に平らになった。周囲のEuclid幾何を超えて、Riemann計量の近似で距離の一致を評価しても、同じ構造と幾何の違いが得られた。動画と文章の表現でもこのパターンが再現された。結果は、関係構造の収束は局所近傍を超えて全体に広がる一方、計量幾何の収束はかなり弱いことを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The Platonic Representation Hypothesis suggests that increasingly capable models converge toward shared representations. Recent work narrows this claim to shared local neighborhood relationships, finding that capacity-dependent trends in several global similarity measures largely disappear after calibration. We challenge this interpretation by showing that prior local-global comparisons confound structural scale (local versus global) with what is compared: relational structure, defined by which samples are related, versus metric geometry, characterized by quantitative relations such as distances, similarities, or correlations. To disentangle these factors, we construct a controlled $2\times2$ framework that evaluates both relational structure and metric geometry at local and global scales. We introduce $H_0$ skeleton overlap as a global counterpart to mutual $k$-nearest neighbors, together with matched distance-aware variants. Across vision-language models, relational structure exhibits robust representational convergence at both scales after calibration, whereas increasingly stringent distance agreement substantially weakens alignment and progressively flattens the capacity-dependent trend. We further extend the analysis beyond ambient Euclidean geometry by evaluating distance agreement under a Riemannian metric approximation and recover the same structure-geometry pattern. The pattern is also reproduced in video-text representations. Together, these results show that relational convergence extends beyond local neighborhoods to global spanning structure, whereas metric geometry exhibits substantially weaker convergence.

著者のコメント

33 pages, 12 figures, 6 tables

arXiv ID: 2609.27252 / 要約の誤りについて