arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

白色化した埋め込みのノルムが測っているもの

Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures

Mohammed Ahnouch, Lotfi Elaachak

この論文をやさしく読む

ひとことで言うと

白色化した埋め込みの大きさは尤度よりも、意味的な異例さを測ると示した研究。

何に役立つ?

埋め込みを使う異常検知や確率の較正を評価する際の参考になる。

この研究の面白いところ

白色化が弱いスペクトル方向の雑音を強める仕組みを示した点。

どこまで分かった?

末尾に原入力の余分な文字列があるが、本文にある結論までを訳した。ガウス閾値の誤差は評価した設定での主張である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

基盤モデルの埋め込みを白色化し、そのノルムの二乗を学習不要な尤度の代わりに使う方法は、白色化した各座標が標準正規分布に近く見えることから動機づけられている。本研究は、この現象が射影に関する中心極限定理から従うため、座標を合わせた同時分布がガウス分布であることを意味しないと示す。複数のエンコーダーと三種類の学習目的で、白色化後の半径はガウス分布の基準より系統的に広がっており、平均と共分散を同じにした分布上の複製と比べても同様だった。経験的なノルム統計量と理論値の一致としてよく報告される現象は、同じ標本で白色化することから代数的に生じるもので、ガウス性の証拠にはならないことも示す。仕組みとして、白色化はエンコーダーのスペクトル上の階層を逆転させ、ノルムの二乗への寄与を、主として雑音を表すほぼ縮退した方向へ移す。そこでの変動は主に入力ごとの単一の尺度で決まり、その尺度を二つのモーメントから推定すると、追加の自由パラメータなしに、分離したスペクトルの前半と後半の依存関係を予測できた。したがって白色化ノルムの二乗は対数尤度より、意味的にどれほど異例かを表すマハラノビス尺度と解釈する方がよい。この解釈は実用上の有効性と較正の失敗の両方を説明する。非パラメトリックな密度推定や異なる目的で学習したエンコーダーと整合する形で異例の標本を順位付け・検出できる一方、ガウス分布を仮定した裾の閾値は桁違いに不正確になりうる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Whitening a foundation-model embedding and using its squared norm as a training-free likelihood surrogate is motivated by the observation that whitened coordinates often appear approximately standard normal. We show that this observation follows from the projection central limit theorem and therefore does not imply a Gaussian joint distribution. Across multiple encoders and three training objectives, we find systematic over-dispersion of the whitened radius relative to the Gaussian reference, including against distributional clones with identical mean and covariance. We further show that the commonly reported agreement between empirical and theoretical norm statistics is an algebraic consequence of in-sample whitening and does not constitute evidence for Gaussianity. We identify the mechanism behind this behavior: whitening reverses the encoder's spectral hierarchy, shifting the contribution to the squared norm toward near-degenerate directions that encode predominantly noise. In these directions, the dominant variability is governed by a single input-dependent scale. We estimate this scale from two moments and use it to predict, without additional free parameters, the cross-dependence between disjoint spectral halves. These results indicate that the squared whitened norm is better interpreted as a Mahalanobis measure of semantic atypicality than as a log-likelihood. This interpretation explains both its practical effectiveness and its calibration failures: the statistic can rank and detect atypical samples consistently with nonparametric density estimates and across encoders trained with different objectives, while Gaussian tail thresholds can be inaccurate by orders of magnitude. etc.

著者のコメント

16 pages, 2 figures

arXiv ID: 2609.23117 / 要約の誤りについて