モデルが外に示す中間推論の効率と構成を比較
Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
この論文をやさしく読む
ひとことで言うと
言語モデルが外部に出す中間的な説明を比較し、使うトークン数や推論の組み立て方の違いを分析した研究です。
何に役立つ?
モデルの最終得点だけでなく、表出された解法がどの程度簡潔で、どのように構成されるかを評価する手掛かりになります。
この研究の面白いところ
公開モデルで通常のCoTとの性能比較を行った後に、非公開モデルへ分析を広げています。説明の長さだけでなく、ステップの種類や推論木も比較対象にしています。
どこまで分かった?
要旨自体が、得られた記録は事後的な理由付けかもしれないと認めています。解答性能が近いことだけでは内部の実際の思考を忠実に取得できたとは証明できません。内部処理についての記述は著者の解釈として読む必要があります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
最先端言語モデルの急速な能力向上は、広く推論能力の改善によるものと考えられている。しかし、非公開システムでは生の思考の連鎖(CoT)の記録が隠されているため、これを検証できない。本研究では、標準APIの機能を通じて簡単なカスタムツールを登録し、最先端モデルに中間推論を外部へ表出させる。こうした記録は本来の推論ではなく事後的な理由付けかもしれないため、まず公開モデルで標準のCoTと比較し、その後、GPT-6 Astraを含む非公開の最先端モデルへ対象を広げる。抽出した推論は、競技数学、科学、コード生成の各分野において、標準の推論と同等の性能を持ち、推論なしのベースラインを大きく上回ることが分かった。 続いて、最先端モデルが中間推論をどう構成するかを特徴付ける。トークン効率、推論ステップの種類、誘導した推論木を通じて、推論を外部へ示し、圧縮し、整理する方法に体系的な違いを見いだす。Astraはトークン効率の高い方向付けられた推論を示し、正しい経路をより早く選ぶ一方、初歩的な手順は内部で処理し、重要な推論だけを外部へ示すことが分かった。これらの結果は、ベンチマークの点数を超えて、最先端モデルの推論を行動から捉える視点を与える。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reasoning. Because these traces may reflect post-hoc rationalization rather than genuine reasoning, we first evaluate against native CoT on open-source models and extend to closed-source frontier models including GPT-6 Astra. We find that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines, across competition mathematics, science, and code generation. We then characterize how frontier models structure their intermediate reasoning. Across token efficiency, reasoning-step types, and induced reasoning trees, we identify systematic differences in how models externalize, compress, and organize reasoning. We find that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning. These findings provide a behavioral lens on frontier-model reasoning beyond benchmark scores.
著者のコメント
33 pages,14 figures
arXiv ID: 2609.26637 / 要約の誤りについて