視覚言語モデルの内部表現に見える「整列の錯覚」
The Alignment Illusion in Multimodal Large Language Models
この論文をやさしく読む
ひとことで言うと
視覚と言語の内部表現が似ていても、モデルが画像の内容を実際に使っているとは限らないことを示した研究。
何に役立つ?
マルチモーダルモデルの内部表現を評価する際、整列スコアを課題性能と照らし合わせるのに役立つ。
この研究の面白いところ
13モデルで視覚情報をノイズに替えると正解率は下がったが、四つの標準的な整列尺度では一貫して検出できなかった。主角ギャップという別の指標を提案した。
どこまで分かった?
主角ギャップも、無関係な構造画像の条件では内部幾何と正解率が乖離し得る。要旨は単独の万能な評価指標とはしていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
マルチモーダル大規模言語モデル(MLLM)の各層で見られる視覚と文章の類似性は、言語モデルが視覚内容を共通の表現空間へ徐々に取り込む証拠と広く解釈されている。この解釈は、単一の数値で表した整列スコアが、内容のレベルでのモダリティ間の相互作用を反映するという仮定に基づく。本研究は、視覚情報の流れへ制御された介入を加えて、この仮定を検証する。五つの系列に属し、パラメータ数が5億~720億の13のMLLMで、射影器の出力となる視覚トークンをガウスノイズに置き換えると、課題の正解率は急激に低下した。しかし、標準的な四つの単一値の尺度、CKA、SVCCA、MIR、最大主角の余弦は、壊れた視覚情報と元の情報を一貫して区別できなかった。これを「整列の錯覚」と呼び、原因を言語モデル内の共有経路に求める。異方的なMLPの次元縮小射影が視覚トークンと文章トークンを共通の出力方向へ引き寄せ、重みに由来する整列を作る。この成分は実質的に一次元なので、上位二つの主角の余弦の差として定義する主角ギャップ(PA gap)を導入し、重みに由来する類似性と多方向の視覚的構造を分ける。視覚情報を段階的に壊した場合、PA gapは検討した単一値のスコアより課題の正解率を一貫して追跡した。一方、構造を持つが無関係な画像を与えると、内部の幾何学的な状態と課題正解率が乖離する条件も明らかにした。したがって、MLLM内部の視覚と文章の整列は、内容レベルの相互作用を直接表す代用指標ではなく、言語モデル内の視覚情報の幾何学的な診断として読むべきであり、制御された課題の証拠と照らし合わせて初めて有益になる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families spanning 0.5B to 72B parameters, replacing projector-output visual tokens with Gaussian noise sharply reduces task accuracy, yet four standard scalar measures (CKA, SVCCA, MIR, and the leading principal-angle cosine) fail to consistently separate the corrupted stream from the original. We call this failure the alignment illusion and trace it to the shared language-model pathway: anisotropic MLP down-projections pull visual and text tokens toward common output directions, producing weight-induced alignment. Because this component is essentially one-dimensional, we introduce the principal-angle gap (PA gap), defined as the difference between the top two principal-angle cosines, which separates weight-induced similarity from multi-directional visual structure. Under graded visual corruption, the PA gap tracks task accuracy more consistently than the scalar scores we consider; under a structured but irrelevant image, it further exposes regimes in which internal geometry and task accuracy come apart. Internal visual-text alignment in MLLMs is therefore best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.
著者のコメント
Accepted to NeurIPS 2026
arXiv ID: 2609.30210 / 要約の誤りについて