思考過程の説明が内部の判断に効いているかを介入で調べる
From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness
この論文をやさしく読む
ひとことで言うと
言語モデルの思考過程の文章が、実際の答えを決めた内部概念を反映しているかを、概念の除去実験で調べる研究。
何に役立つ?
思考過程の説明を評価する際、文章がもっともらしいことや内部概念が似ていることに加え、答えへの因果的な寄与を確認するための評価法として役立つ。
この研究の面白いところ
共通の内部概念が多くても、その概念が答えを左右する強さは層によって異なる。重要な概念が説明文に書かれていない場合もある。
どこまで分かった?
結果は五つの言語モデルと四つのデータセット、および共通の疎なオートエンコーダーによる概念近似と除去操作に基づく。説明の忠実性をすべてのモデルで確定する万能な指標とは述べていない。
v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
言語モデルが示す思考過程の説明(CoT)はもっともらしく見えても、モデルの実際の推論を忠実に表していない場合がある。先行研究の多くは入出力の振る舞いや入力への帰属から忠実性を調べ、内部計算は十分に探っていない。本研究では、忠実性を内部概念への根付き方として捉える。言語モデルのCoTによる推論は、説明を使わない直接予測を支える内部概念と同じものを使っているか、また共通の概念が答えを因果的に決めているかを問う。直接予測とCoTを伴う予測の両方を、共通の疎なオートエンコーダーで符号化し、潜在概念を直接比較可能にする。概念レベルの対応について三つの相関指標を導入し、共通概念を除去したときの答えの確率の低下を測る因果指標Δpも導入する。 五つの言語モデルと四つのデータセットでは、相関指標で見た概念の対応はおおむね高かった。しかし、相関指標が分かるのはどの概念が共有されるかであり、答えへの因果的な寄与の大きさではない。Δpで見る因果的な忠実性はモデルの層の深さによって大きく変わり、最後の層ではなく中盤から後半の層で最大になった。モデルの規模によって層ごとの形も変わった。また、因果的に重要な共通概念がCoTの文章に必ずしも表現されるわけではなかった。これらの違いは、表面上の説明や内部表現の対応だけで忠実性を確実に判断できず、CoTに関わる内部概念が実際に予測を動かすかを介入で検査する必要があることを示唆する。
v2の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-23 · v2
- 査読・掲載
- 査読状況未確認
更新履歴
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large language model's (LLM) CoT reasoning engage the same internal concepts that support the LLM's direct prediction, and do the shared concepts causally drive its answer? Encoding a prediction pass and a CoT pass with a single shared sparse autoencoder (SAE), a reliable approximator of the latent concepts LLMs use, makes their internal concepts directly comparable. We introduce three correlational metrics of concept-level alignment and a causal metric, $\Delta p$, which ablates the shared concepts and measures the drop in answer probability. Across five LLMs and four datasets, concept alignment is generally high, as indicated by the correlational metrics; yet these only identify which concepts are shared, not how much they causally contribute. $\Delta p$ fills this gap: causal faithfulness varies substantially with model depth, peaking at mid-to-late layers rather than the final ones, and model scale reshapes the layer-wise profile. Moreover, causally important shared concepts are not always verbalized in the CoT. These dissociations suggest that faithfulness cannot be reliably assessed from surface-level or representational correspondence alone; assessing it requires causal tests of whether the internal concepts underlying a CoT actually drive the model's prediction.
著者のコメント
In submission
arXiv ID: 2609.23065 / 要約の誤りについて