AIが書いた推論手順は答えの生成に実際に効いているか
Are Stated Reasoning Steps Causally Load-Bearing?
この論文をやさしく読む
ひとことで言うと
AIが文章で示す推論手順が、答えを作る内部計算に本当に関わるかを活性化の介入で調べた研究。
何に役立つ?
思考の連鎖を監視や評価に使う際、文章の編集だけで忠実性を測る方法の限界を見極める助けになる。
この研究の面白いところ
中間手順の活性化を別の実行と差し替え、予測した特定の答えに切り替わるかを確認する。
どこまで分かった?
2~6段階の合成課題とQwen3の2モデルでの結果。複雑な実世界課題やほかのモデルにも同じ割合が当てはまるとは要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
思考の連鎖(CoT)の監視は、モデルが書いた推論が、答えを直接作る計算を反映すると仮定する。従来の忠実性指標は、推論文を編集して答えの変化を見る、主に行動上の試験だった。本研究は、モデル自身が生成した推論について、活性化のレベルで因果的な忠実性を測ることを目指す。以前の因果的な監査のように性能の低下だけを見るのではなく、介入後にどの答えへ変わるべきかをあらかじめ定める。各介入により、構成上決められた別の対象へ答えが切り替わるかを調べる。 具体的には、2~6段階の合成的な多段階参照課題を使い、モデルが各中間手順を述べるトークン位置で、残差ストリームの活性化を反事実的な実行から得た対応する活性化に差し替える。Qwen3-4Bでは、反応が最も大きいネットワーク中層で、述べられた手順の76.9%±2.8%が答えの生成に因果的に必要だった。無関係な位置で差し替える対照は11.3%、入力文の元の事実を差し替える場合は83%であり、述べられた手順は達成可能な効果の約96%を担った。同じ問題に対する標準的な行動試験では88.2%となり、因果的な忠実性を11.4ポイント過大評価した。項目を対応させた比較で不一致は111対14、p値は10のマイナス15乗未満で、最も易しい問題では差が最大20ポイントだった。Qwen3-1.7Bの因果的忠実性は全体で54.8%と低く、推論が2段階から6段階に深まると68%から30%へ落ちたのに対し、Qwen3-4Bでは比較的変わらなかった。書かれた推論は因果的な意味を持ち得るが、標準的な行動試験はその忠実性を、特に流暢に推論する易しい問題で過大評価しやすい。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Chain-of-thought (CoT) monitoring assumes that the reasoning a model writes reflects the computation that directly produces its answer. Previous faithfulness metrics have been predominantly behavioral, as they simply edit the reasoning text and observe the resulting answer. However, our methodology aims to measure faithfulness causally at the activation level, specifically on self-generated reasoning. Unlike previous causal audits, which measure degradation, our interventions carry a known predicted target. In this way, each patch should switch the answer to a specific counterfactual entity derivable by construction. Specifically, we use synthetic multi-hop lookup tasks (2-6 hops). We patch the residual stream at the token span where the model states each intermediate step with the corresponding activations from a counterfactual run. For Qwen3-4B, 76.9% +/- 2.8% of stated steps are causally load-bearing (CLB) at the most responsive mid-network layer (random-position null: 11.3%; patching the underlying prompt fact: 83%, so stated steps carry approximately 96% of the achievable effect). Moreover, the standard behavioral test on the same items yields 88.2%, which overstates causal faithfulness by 11.4 percentage points (item-matched; 111:14 discordant pairs, p < 1e-15) and, for the easiest items, by up to 20 percentage points. This gap also has a clear capability dimension. Qwen3-1.7B is far less causally faithful overall (54.8%), with its faithfulness collapsing as reasoning depth increases (68% at 2 hops to 30% at 6), while Qwen3-4B remains relatively flat. Although stated reasoning can be causally meaningful, standard behavioral tests tend to overestimate its causal faithfulness, particularly on easier examples where model reasoning appears most fluent.
著者のコメント
NeurIPS 2026, Interpretability as a Science
arXiv ID: 2609.27038 / 要約の誤りについて