文字認識の注意ヘッドから画像の意味を言葉にする
Using OCR Heads to Verbalize Image Semantics
この論文をやさしく読む
ひとことで言うと
画像内の文字を読む仕組みを追うと、鳥の翼など文字でない画像内容も言葉に対応付ける内部の働きが見つかった。
何に役立つ?
視覚言語モデルの隠れ状態が何を表しているかを、語彙によるラベルとして調べる解釈手法に役立つ。
この研究の面白いところ
OCR用と考えられる注意ヘッドが汎用的な意味表現にも関わり、逆変換による概念編集でその役割を因果的にも調べている。
どこまで分かった?
四つのモデルでの解析であり、あらゆる視覚言語モデルで同じ仕組みが成立するとまでは示していない。初期層の言語との整合も、画像内容を常に正確に理解できるという保証ではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚言語モデル(VLM)は、画素をどのように意味へ対応付けるのか。この一般的な問いを理解するため、VLMが光学文字認識(OCR)をどう行うかという狭い問いに注目する。四つのモデルで、OCRに因果的に必要な注意ヘッドを特定し、それらが実際には、あらゆる画像トークンについて解釈可能な意味特徴を出力する汎用的なヘッドであることを発見した。例えば、これらのヘッドを「bike」という語を含む画像トークンへ向けると、Qwen3-VL-8Bは「bike」を出力するが、鳥の翼へ向けると「feathers」というトークンを出力する。 これらのヘッドの注意重みを一つの言語化レンズ変換へまとめ、すべての層の隠れ状態にある解釈可能な意味特徴を明らかにする。語彙空間への射影と組み合わせると、第0層から解釈可能なラベルを得られ、画像表現が実際には初期層から言語と整合していることが分かる。また、この変換の逆を使って、自然な画像中のトラクターをリボルバーへ置き換えるといった、単語ではない概念の編集もできることが分かった。これは、この部分空間がOCR以外にも役立つという因果的な証拠を与える。本結果は、特定の仕組みの研究が、より広い解釈可能性の問題を明らかにする一例である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention heads causally necessary for OCR, and discover that these are in fact general-purpose heads that output interpretable semantic features across all image tokens. For example, pointing these heads at an image token containing the word "bike" causes Qwen3-VL-8B to output "bike," but pointing them at a bird wing causes the model to output the token "feathers." We collapse these heads' attention weights into a single verbalization lens transformation that reveals interpretable semantic features in hidden states across all layers. When combined with projection to vocabulary space, we can obtain interpretable labels starting from layer 0, showing that image representations are in fact aligned with language in early layers. We find that we can also use the inverse of this transformation to edit non-word concepts, e.g., replacing a tractor with a revolver in a naturalistic image, providing causal evidence that this subspace is useful for more than just OCR. Our results are an example of how the study of specific mechanisms can shed light on broader interpretability problems.
著者のコメント
21 pages, 22 figures
arXiv ID: 2609.18823 / 要約の誤りについて