arXiv論文メモ
新着一覧
cs.CV / cs.LG · 査読状況未確認

視覚モデルの局所配置学習は因果的な回路を集中させる

Topographic Training Concentrates Causal Circuits Without Improving Neuron Monosemanticity

Gautam Ranka, Shubham Santosh Pandere, Aiden Dsouza

この論文をやさしく読む

ひとことで言うと

似た役割を持つ計算をモデル内で近くに集めるよう学習すると、予測に影響する回路が局所的にまとまりました。ただし、個々のニューロンが一つの意味だけを表すようになったわけではありません。

何に役立つ?

モデルの内部を調べる際、ニューロン単体の分かりやすさだけでなく、回路のまとまりも評価する必要があることを示します。解釈しやすい視覚モデルの学習方法を設計する参考になります。

この研究の面白いところ

同じサイズの無作為なユニット集合との比較で、因果的十分性が2.79倍になっています。一方で単一意味性の指標は変わらず、二つの「分かりやすさ」が一致しないことを捉えています。

どこまで分かった?

実験はImageNet-100で学習したViTに基づきます。SAEの不活性特徴の増加を含む変化が報告されており、すべての解釈指標が一様に改善したという結果ではありません。要旨には他のモデル系列への検証はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Vision Transformerの機構的解釈可能性の研究は、モデルの計算を人間が読める単位へ分解することを目指すが、学習された表現では、各ニューロンに多くの概念が絡み合っている。特徴の重ね合わせは、この分解を妨げる中心的な障害と広くみなされている。しかし、スパースオートエンコーダーや辞書学習など、その軽減策の多くは学習後に適用され、元のネットワークを変更しない。本研究では、空間的な局所性を促す学習損失TopoLossが、標準的な機構解釈ツールでの解釈可能性を高める、軽量な学習時の事前制約として働くかを問う。 TopoLossの重みαを複数設定してImageNet-100上でViTを学習し、活性のパッチングによって空間配置に基づくクラスターの因果的十分性を測る。同じ残差ストリームに適合させたスパースオートエンコーダー(SAE)を用い、特徴の幾何学的構造も測定する。α=1.0では、空間配置に基づくクラスターの因果的十分性は、同じ大きさの無作為なユニット集合の2.79倍となり、その効果はαとともに単調に増加する。SAEのL0スパース性指標は11%低下し、不活性特徴の割合は19倍になる一方、ニューロン単位の標準的な単一意味性スコアは変化しない。 これは、空間配置への圧力が回路レベルで働き、個々のニューロンの概念を解きほぐすことなく、因果的な寄与を空間的に局所な構造へ集中させることを示している。この乖離は、現在のニューロン単位の単一意味性指標が、実際に存在するある種の解釈可能性の向上を捉えられないことを示唆する。また、低コストな構造上の事前制約を、学習後のツールを補完する有効な学習時の方法として位置づける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposition is widely treated as the central obstacle to this decomposition, yet most mitigations (sparse autoencoders, dictionary learning) are post-hoc and leave the underlying network unchanged. We ask whether a spatial-locality training loss (TopoLoss) can act as a lightweight, training-time prior that improves interpretability of standard mech-interp tools. Training ViT on ImageNet-100 across multiple TopoLoss weights $\alpha$, we measure causal sufficiency of topographic clusters via activation patching and feature geometry via sparse autoencoders fit to the same residual stream. At $\alpha=1.0$, topographic clusters are 2.79$\times$ more causally sufficient than random unit sets of the same size, with the effect increasing monotonically in $\alpha$. SAE L0 sparsity decreases by 11% and dead-feature fraction rises 19-fold, yet standard neuron-level monosemanticity scores are unchanged, indicating that topographic pressure acts at circuit level, concentrating causal mass into spatially local structures without disentangling individual neurons. This dissociation suggests current neuron-level monosemanticity metrics are insensitive to a class of real interpretability gains, and positions cheap architectural priors as a viable training-time complement to post-hoc tooling.

著者のコメント

Accepted at the Mechanistic Interpretability Workshop at ICML 2026

arXiv ID: 2609.24379 / 要約の誤りについて