大規模言語モデルで文法専用ニューロンはまれ
Grammatical "grandmother neurons" are rare in LLMs
この論文をやさしく読む
ひとことで言うと
LLMの文法知識が特定のニューロンだけに入っているのかを調べ、強く専門化したものはまれだと報告します。
何に役立つ?
LLM内部の言語表現を調べる際、追加の分類器を学習させずに単一ニューロンの反応を評価する方法として役立ちます。
この研究の面白いところ
言語学的な最小対を使い、68種類の現象と7つのチェックポイントを比較しました。単一ユニットの反応と、モデルの実際の言語能力が大きく異なる点を示しています。
どこまで分かった?
単一ニューロンが文法と非文法を区別しても、そのニューロンがモデルの判断に必要とは限りません。結果の範囲は要旨に記載された言語現象とチェックポイントです。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)が言語の構造をどのように符号化しているかは、解釈可能性研究の基本的な課題である。その調査には診断用の分類器、いわゆるプローブが広く使われるが、補助分類器を学習させると、その学習能力や較正の問題が混ざり、モデル自体の表現とプローブが課題を学んだ能力を区別しにくくなる。著者らはこの制約に対処するため、個々のニューロンが示す言語的な選択性を、プローブなしで特定する枠組みを提案する。言語学的な最小対の統制された比較を利用し、パラメータを更新せずに、単一ニューロンが文法的な構文と非文法的な構文をどの程度安定して区別できるかを直接測るニューロン分離指数(NSI)を導入する。 68種類の言語現象と7つのチェックポイントにNSIを適用した結果、主に三つの傾向が見られた。第一に、形態と統語の区別は、統語と意味の接点や概念的な区別よりも早い段階で、補正前の分離能力がほぼ最大に達した。第二に、データの並べ替えによる正規化の後では、単一ユニットの選択性はまばらで弱く、対象が狭かった。平均的な言語現象に反応するユニットは少数で、強い選択性を持つ「おばあさんニューロン」はまれだった。第三に、表現ベクトル全体の線形分離可能性、単一ニューロンの選択性、モデルの振る舞いとしての言語能力は大きく分かれていた。特定のニューロンを除去する実験でも、活性化の選択性と、そのニューロンへの因果的な依存は区別された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Understanding how Large Language Models (LLMs) encode linguistic structures remains a fundamental challenge in interpretability research. While diagnostic classifiers (or "probes") are widely used for this task, they face significant methodological criticism: training auxiliary classifiers introduces capacity confounds and calibration issues, often making it difficult to distinguish the model's intrinsic representations from the probe's ability to learn the task. To address these limitations, we introduce a probe-free framework for localizing linguistic selectivity at the individual neuron level. Leveraging the controlled contrasts of linguistic minimal pairs, we propose a Neuron Separability Index (NSI), a metric that directly quantifies how reliably single neurons differentiate grammatical from ungrammatical constructions without parameter updates. Applying NSI across 68 linguistic paradigms and seven checkpoints reveals three main patterns: 1) raw separability reaches near-peak levels earlier for morphological and syntactic distinctions than for syntax-semantics interface and conceptual distinctions. 2) after permutation normalization, single-unit selectivity is sparse, weak, and narrowly tuned: only a small fraction of units are sensitive to an average paradigm, and strongly selective "grandmother neurons" are rare. 3) whole-vector linear separability, single-neuron selectivity, and behavioral competence are largely dissociated, and targeted ablations further separate activation selectivity from causal reliance.
著者のコメント
Accepted at COLM 2026. 28 pages
arXiv ID: 2609.29328 / 要約の誤りについて