幻覚を示すニューロンは一意に特定できるか
Hallucination Neurons and Where to Find Them: An Investigation into the existence of Hallucination Neurons
この論文をやさしく読む
ひとことで言うと
幻覚に関係するニューロンを検出できても、特定のニューロンだけが担うとは言えないことを調べた研究です。
何に役立つ?
言語モデルの内部を監査する研究で、検出性能と原因の局在を分けて評価する手順として役立ちます。
この研究の面白いところ
検出と介入の効果を再確認しつつ、選ばれたニューロンの多くが他の特徴と強く相関することを示しました。
どこまで分かった?
評価対象は要旨に挙げたモデルと3データセットです。ニューロンの選択が一意でないため、少数の特定位置に機能を帰属させる解釈には限界があります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルの解釈可能性研究では、少数のニューロンが事実の想起、安全性への整合、幻覚などを検出し、因果的に左右すると主張する疎なプローブがよく使われる。しかし、相関の強い高次元特徴にL1正則化を用いた場合の既知の問題に照らした検証は少ない。著者らは、特徴間の相関、ブートストラップでの安定性、疎な順位と密な順位の不一致、介入の比較基準、データセットをまたぐ評価からなる5段階の診断手順を、疎なニューロンの局在を主張する際の最低限の基準として提案する。 提案手順を使い、TriviaQA、BioASQ、NQ-Openとオープンソースの言語モデルで、先行研究の幻覚関連ニューロンを調べた。検出結果はモデルとデータセットをまたいで再現され、TriviaQAとBioASQでは元の研究が報告したAUROCの差を上回った。対応するデータセットではGemma 3 4BがMedGemma 4Bより一貫して高く、差はTriviaQAで+0.311対+0.235、BioASQで+0.474対+0.455、NQ-Openで+0.128対+0.112だった。500例、乱数種5通りでの因果検証では、同じ層から無作為に選んだ比較基準を超える統計的に有意な効果があった。一方、選ばれたニューロンの位置は一意ではない。Gemma 3 4Bの3設定で選んだ22個のうち19個は他の特徴とピアソン相関係数の絶対値が0.7を超え、ブートストラップでの選択安定性は中程度、疎な順位と密な順位の重なりは小さかった。疎な予測構造とニューロン選択の非一意性は両立するため、検出の主張と局在の主張を分けるには定型的な診断が必要である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Interpretable machine learning for Large Language Models (LLMs) increasingly relies on sparse probing methods that identify small sets of neurons claimed to detect and causally influence behaviors such as factuality recall, safety alignment, and hallucination. These claims have important implications for model auditing and behavioral steering, yet they are rarely tested against known failure modes of $L_1$-regularized probing in correlated, high-dimensional feature spaces. We propose a five-step diagnostic protocol covering feature correlation, bootstrap stability, sparse versus dense ranking disagreement, intervention baselines, and cross-dataset evaluation as a minimum standard for sparse-neuron localization claims. We investigate prior work using our proposed approach, specifically on H-neurons using open-source LLMs across TriviaQA, BioASQ, and NQ-Open datasets. Our results demonstrate detection replicates across both models and datasets, and exceeds the original reported AUROC gaps for TriviaQA and BioASQ datasets. Gemma 3 4B consistently outperforms MedGemma 4B on matched datasets, with AUROC gaps of +0.311 versus +0.235 on TriviaQA, +0.474 versus +0.455 on BioASQ, and +0.128 versus +0.112 on NQ-Open respectively. Causal validation at $n = 500$ with five random seeds shows statistically significant effects beyond random same-layer baselines. At the same time, the diagnostic results indicate that the selected neurons are not uniquely localized. Across the three Gemma 3 4B settings, 19 of 22 selected H-Neurons have Pearson $|r| > 0.7$ with other features, bootstrap selections show only moderate stability, and sparse and dense rankings overlap only weakly. Our findings show that sparse predictive structure can coexist with non-unique neuron selection. Routine diagnostic validation is necessary to distinguish detection claims from localization claims in mechanistic interpretability.
arXiv ID: 2609.29781 / 要約の誤りについて