arXiv論文メモ
新着一覧
cs.LG / cs.CL · 査読状況未確認

言語モデルの自己修復に見える応答を固定の作用で説明

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu

この論文をやさしく読む

ひとことで言うと

一部を取り除くとモデルが自分で直したように見える現象を、もともと存在する固定の応答で説明しようとする研究です。介入の強さを連続的な軸として扱います。

何に役立つ?

モデルの構成要素を除去して役割を調べる実験の解釈に役立ちます。除去後の補償を新しい適応とみなす前に、既存の重みの作用で説明できるかを検討できます。

この研究の面白いところ

介入の種類を単なる有無ではなく強度λで整理し、応答の傾きから対抗・増強を判定します。その傾きの大きさを固定重みから予測する点が特徴です。

どこまで分かった?

法則に従ったのは4モデルの対象81方向中68方向、GPT-2 Smallの対象10ヘッド中7個です。全構成要素で成立したわけではありません。自己修復一般の完全な解明ではなく、記載された課題・回路での説明と証拠です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

言語モデルの構成要素を除去すると、他の構成要素が調整して補償しているように見えることが多い。自己修復と呼ばれるこの現象は繰り返し観測されてきたが、その仕組みは明らかでない。これまでで最も体系的な研究は、自己修復には雑音が多く、単一の説明はありそうにないと結論した。私たちは、除去前から存在するゲインという、一つの説明があると主張する。因果的に重要な構成要素への介入はどれも、反実仮想的な対比の符号付き強度λという座標軸上の一点とみなせる。したがって、従来の除去法はこの軸上の未較正の点である。 細かく分けた単位rの因果的な修復応答が、アフィン則E_r(λ)=own_r+γ_r λに従うことを示す。傾きγ_rは、除去の有無にかかわらず一貫してモデルに影響する固定係数であり、その符号が、取り除いた信号に対抗するか強めるかを決める。Gemma、Qwen、LLaMA、Mistralという異なる系列の4モデルを使う事実判定課題で、MLPニューロン、OVニューロン、特異方向を含む構成要素がこのアフィン則に従うことを確認した。下流方向81個のうち、合計68個が該当した。さらに、固定された重みからγ_rの大きさを予測できる。 GPT-2 SmallのIOI回路では、介入の影響が届く10個のヘッド中7個がこの法則に従い、その7個はすべて対抗作用を持つ。この見方では、自己修復に見えるものは、中心に対比信号が現れたときに、対抗作用を持つ要素が通常どおりの動作をしていることになる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Ablate a component of a language model, and other components often appear to adjust and compensate. This phenomenon, termed self-repair, has been observed repeatedly, but its mechanism remains unclear. The most systematic study to date concluded that self-repair is noisy and unlikely to have a single explanation. We argue that it has one: a gain already present before any ablation. Any intervention on a causally important component can be viewed as a point on a coordinate axis $\lambda$, the signed strength of a counterfactual contrast. Hence, conventional ablation methods are uncalibrated points on this axis. We show that the causal repair response for a fine-grained unit $r$ is governed by an affine law, $E_r(\lambda)=\mathrm{own}_r+\gamma_r\lambda$. The slope $\gamma_r$ is a fixed coefficient that consistently influences the model, with or without ablation, and its sign determines whether the unit counteracts or reinforces the removed signal. On a factual-verdict task across four models from distinct families (Gemma, Qwen, LLaMA, and Mistral), we identify components including MLP neurons, OV neurons, and singular directions that follow this affine law, 68 of 81 downstream directions in all. Moreover, we can anticipate the magnitude of $\gamma_r$ from the fixed weights. On the IOI circuit of GPT-2 Small, seven of the ten heads the intervention can reach follow the law, and all seven are counterweights. From this perspective, what may appear as self-repair is a counterweight performing its usual operation when the contrastive signal emerges at the core.

arXiv ID: 2610.02173 / 要約の誤りについて