思考連鎖のエントロピーは信頼性の手掛かりになるか
Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction
この論文をやさしく読む
ひとことで言うと
LLMの推論中の不確かさの変わり方が、答えの正しさの手がかりになるという先行結果を、事前登録して再検証しています。
何に役立つ?
推論の信頼性指標を選ぶ際、軌跡の形と不確かさの総減少量を分けて評価するために役立ちます。
この研究の面白いところ
GSM8KとMATH-500の全テストを4モデルで調べ、形の信号は再現した一方、総減少量は設定で結果が分かれました。最終段階のエントロピーだけの方が良い場合も探索的に示します。
どこまで分かった?
推論蒸留モデルでは二値の形信号が約100件に1件しか出ず、登録した比較を推定できませんでした。7つの手順差や未記載の積分範囲の影響もあり、万能な判定指標ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
本研究は、Zhaoが2026年に報告した乖離を独立に再現する経験的研究である。大規模言語モデルの思考連鎖におけるエントロピー軌跡の形状は最終回答が正しいかどうかを予測する一方、エントロピーが全体としてどれだけ低下したかという大きさは予測しない。この乖離を再現する価値があるのは、大きさに関する結果が、1つのモデルを1つの乱数シードで300問だけ実行した結果に基づくのに対し、形状に関する結果は両方のベンチマークと別のモデル系列で、全規模で報告されているためである。 確認実験を行う前にOSFへ登録したうえで、元研究が検証していない種類の推論蒸留モデルを1つ含む、4つのオープンウェイトモデルを用い、GSM8KとMATH-500のテストセット全体で再現実験を行った。形状のシグナルは再現された。大きさのシグナルは設定によって結果が分かれた。基準モデルでは、単調な思考連鎖と非単調な思考連鎖の正答率の差はGSM8Kで9.6ポイント、MATH-500で27.5ポイントだった。一方、全エントロピー低下量と正しさの順位相関は、GSM8Kで-0.018、MATH-500で+0.414だった。 推論蒸留モデルでは、形状シグナルの二値版が発火するのはおよそ100連鎖に1つであり、登録した差を推定するには少なすぎたが、違反回数を段階的に数える指標はそこでなお予測性を示した。探索的比較では、最終ステップのエントロピーだけを用いる方法が、8つのモデルとベンチマークの組み合わせすべてでROC面積により二値形状フラグを上回り、元研究が報告するリスク・カバレッジ面積でも、元研究が積分範囲を明記していないため範囲によって異なるものの、6または7組で上回った。 本研究は、7つの文書化されたプロトコル差がある条件でテストセット全体の規模で形状シグナルを再現し、大きさのシグナルが成立する設定と成立しない設定を整理し、元研究が報告していない4つのプロトコル依存性を測定した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
This empirical study is an independent reproduction of the dissociation Zhao reported in 2026. The shape of a large language model's chain-of-thought entropy trajectory predicts whether the final answer is correct, while the magnitude of its total entropy drop does not. The dissociation merits reproduction because the magnitude half rests on a single 300-problem run with one model at one seed, while the shape half was reported at full scale on both benchmarks and on a second model family. Registered at OSF before any confirmatory run, the reproduction crosses the complete GSM8K and MATH-500 benchmark test sets with four open-weight models including one reasoning-distilled model of a kind the original did not test. The shape signal replicates. The magnitude signal divides by setting. On the anchor model the accuracy gap between monotone and non-monotone chains is +9.6 percentage points on GSM8K and +27.5 on MATH-500, while the rank correlation of the total entropy drop with correctness is -0.018 on GSM8K and +0.414 on MATH-500. On the reasoning-distilled model the binary form of the shape signal fires on about one chain in a hundred, too few to estimate the registered contrast, while the graded violation count remains predictive there. In an exploratory comparison the final-step entropy alone outperforms the binary shape flag in all eight model-by-benchmark cells by ROC area, and in six or seven by the risk-coverage area the original reports, depending on an integration range the original does not state. The study contributes a reproduction of the shape signal at full test-set scale under seven documented protocol differences, a map of the settings where the magnitude signal holds and fails, and measurements of four protocol dependencies the original does not report.
arXiv ID: 2609.19606 / 要約の誤りについて