AIの内部回路を選ぶ評価指標が誤順位を生む
Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability
この論文をやさしく読む
ひとことで言うと
AIの内部計算を説明する回路を探す際、評価指標そのものが、元のモデルを再現しにくい回路を高く評価することがあると示しています。
何に役立つ?
内部回路の発見手法を比較するとき、探索性能だけでなく評価指標の妥当性を確認するために役立ちます。
この研究の面白いところ
探索アルゴリズムを使わない制御編集でも問題を確認しています。回路自体を変えず、介入で置き換えた信号の一部を戻すだけで100件中96件の誤順位を直した点が、評価時の文脈の影響を示しています。
どこまで分かった?
評価は4つの人手参照課題とInterpBenchに基づきます。元モデルとの一致は正しい回答との一致ではなく、通常は元モデルの間違いも再現する基準です。すべての忠実度指標に同じ誤順位率が当てはまるとは述べられていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
機構的解釈可能性は、モデルの振る舞いを生む内部計算の復元を目指す。自動回路発見の進歩は、より良い寄与度評価や最適化によって、より良い機構を特定できるという探索問題として語られることが多い。これは、より良い回路が見つかったとき、評価目的がそれを認識できると仮定している。本研究では、介入によって定義された忠実度が、同じ大きさでありながらモデルの振る舞いをより悪く再現する回路を好む場合があり、評価目的そのものの段階で復元のずれが生じることを示す。 人が参照回路を作った4課題とInterpBenchを用い、通常の再サンプリングを固定した条件のもとで、検証用の忠実度と、評価用に取り置いたプロンプト上の振る舞いを比較する。振る舞いの基準は、Greater-Than課題では意味上の正確さを用い、それ以外ではモデルの誤りも含めた、介入前の完全なモデルとの一致とする。参照回路に対する制御された編集によって、発見アルゴリズムがなくても誤った順位付けが生じることが分かり、EAP、EAP-IG、ACDC、Edge-SPの出力でも同じ問題が見られる。再サンプリング条件では、人の参照回路がある課題において、これらの手法の候補対の9.4%〜41.2%をKLが誤って順位付けする。 その説明として、文脈の歪みを調べる。除外した信号を置き換えると、残した構成要素が動作する際の入力が変わる。介入先の入力について完全なモデルを実行したときの信号を選択的に戻すことで、発見候補集合における持続的なKL誤順位100件のうち96件が、検証用・取り置きの両方のプロンプトで修復される。回路と元の振る舞いのスコアは変わらない。これらの知見は、評価目的が誤った候補に報酬を与える場合、発見手法の改善だけでは不十分である理由を示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Mechanistic interpretability aims to recover the internal computations responsible for model behavior. Progress in automated circuit discovery is often framed as a search problem: better attribution or optimization should identify better mechanisms. This assumes that the evaluation objective can recognize a better circuit once it is found. We show that intervention-defined faithfulness can instead prefer an equally sized circuit that reproduces the model's behavior less well, creating an objective-level recovery gap. Across four human-reference tasks and InterpBench, we compare validation faithfulness with behavior on held-out prompts under fixed ordinary resampling. The behavioral criterion is agreement with the intact model, including its mistakes, except on Greater-Than, where we use semantic accuracy. Controlled reference edits reveal misranking without any discovery algorithm, and outputs of EAP, EAP-IG, ACDC, and Edge-SP exhibit the same failure. Under resampling, KL misranks 9.4%-41.2% of candidate pairs across these methods on the human-reference tasks. We investigate context distortion as an explanation: replacing excluded signals changes the inputs on which retained components operate. Restoring selected signals from the recipient's intact-model execution repairs 96 of 100 persistent KL misrankings from the discovery pool on both validation and held-out prompts. The circuits and their original behavioral scores remain unchanged. These findings show why better discovery alone is insufficient when its objective rewards the wrong candidate.
著者のコメント
34 pages, 2 figures
arXiv ID: 2610.02098 / 要約の誤りについて