言語モデルの数学的推論は題材より解法でまとまる
Math Reasoning in LLMs is Organized by Approach, Not Topic
この論文をやさしく読む
ひとことで言うと
数学を解く言語モデルの内部表現が、問題の題材より、使う解法に沿ってまとまる証拠を示した。
何に役立つ?
数学的推論の評価問題や学習用データを設計するとき、題材だけでなく解法の種類も確認する助けになる。解法別の設計で性能が向上したことを実証したものではない。
この研究の面白いところ
生成した解答を再生して内部活性化を分析し、教師なしクラスタリング、別モデルの判定、解法を指定する介入で構造を調べた。
どこまで分かった?
対象は公開の8モデルと5種類の数学的推論の出典である。解法の一貫性の一部は別の言語モデルによる判定に基づく。
v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
数学的推論のベンチマークは通常、題材ごとに整理される。しかし言語モデルは内部の計算を、題材別の技能ではなく、再利用できる推論の進め方によって整理している可能性がある。本研究は、数学問題を解ける公開言語モデルの内部表現が、題材ごとの下位技能と解法のどちらに従って構成されるかを調べ、解法が重要である証拠を提示する。 まずモデルに解答を生成させ、次に入力文と生成文の軌跡をそのまま再生して、推論中のトークンについて活性化の重要度を表す特徴を取り出す「生成・再生」手順を導入する。8モデルと5種類の数学的推論の出典について、この特徴を教師なしでクラスタリングし、得られた構造を、構造的・意味的な検査と介入試験で評価する。モデルと出典を組み合わせた40条件すべてで、得られたクラスターは大きさをそろえた無作為の基準を上回った。独立した2つの先端言語モデルによる判定では、実際のクラスターの77~82%に解法レベルの一貫性が見られたのに対し、同じ出典内の対照群では6~11%だった。題材が一つにそろったクラスターでも、その題材より細かいラベルが付くことが多かった。解法を指定する入力では、求める解法を変えると8モデル中7モデルの条件で所属クラスターが変わった一方、言い換えではおおむね保たれた。これらの結果は、数学を解ける言語モデルの内部計算がベンチマークの題材より解法によって整理されることを示す。題材ごとに層化したベンチマークや題材を均等にした学習用文章でも、重要な解法の偏りを見落としうる。
v2の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-24 · v2
- 査読・掲載
- 査読状況未確認
更新履歴
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and we present evidence that the approach is the key. We introduce a generation-replay protocol: a model first generates a solution, after which we replay the exact prompt-plus-generation trajectory and extract activation-importance signatures over the reasoning tokens. We cluster these signatures without supervision across eight models and five mathematical reasoning sources, then evaluate the recovered structure with structural, semantic, and intervention tests. Across all 40 model-source cells, the recovered clusters outperform matched-size random baselines. Two independent frontier-LLM judges find approach-level coherence in 77-82% of real clusters versus 6-11% in within-source controls, and topic-pure clusters usually receive labels finer than the topic itself. In approach-controlled prompting, changing the requested reasoning approach shifts cluster assignment in seven of eight model conditions, whereas paraphrases largely preserve it. These results indicate that math-capable LLMs organize internal mathematical computation by reasoning approach rather than benchmark topic. The implication is that topic-stratified benchmarks and topic-balanced training corpora can still miss the axis that matters: even deliberately topic-balanced corpora may remain imbalanced over reasoning approaches.
arXiv ID: 2609.27041 / 要約の誤りについて