数学推論で必要な基本要素の発見能力を診断・改善
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
この論文をやさしく読む
ひとことで言うと
数学問題を正解できるかだけでなく、解くための基本要素を見つけ、作り、理解し、使う能力を分けて調べる研究です。特に必要な要素を発見する段階に注目しています。
何に役立つ?
数学推論モデルがどこでつまずくかを診断し、事後学習の対象を選ぶのに役立ちます。基本要素に導かれた推論を自己蒸留で取り込む改善法も評価しています。
この研究の面白いところ
基本要素が与えられると実行できても、自分ではそれを発見できないという能力差を捉えます。正答率だけでは見えにくい弱点を、学習による修復へつなげています。
どこまで分かった?
要旨のベンチマーク名と手法名は未展開のLaTeX記号であり、正式名称を読み取れません。モデル規模をまたぐ改善は報告されていますが、具体的なモデル、改善量、4能力の詳細な評価手順は要旨にありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)は高度な数学の問題で目覚ましい能力を示しているが、その解答の基礎となる構造的な数学理解を備えているかは、依然として明らかでない。本論文では、異なる能力を診断することから、その知見を事後学習の改善に利用することまで、LLMの数学理解を体系的に研究する第一歩を踏み出す。 第一に、構造的な数学理解を調べるため、Mathematical Primitive(数学的な基本要素)という概念を導入する。そして、発見、生成、理解・消化、実行という4つの異なる側面から数学推論を評価する、新しいベンチマーク[原文表記:\hlei{}]を提案する。第二に、体系的な診断から、解答の正確さだけでは異なる能力の構成が見えなくなること、基本要素を与えると潜在的な実行能力を大きく引き出せること、そして数学推論の主なボトルネックが発見であることを示す。事後学習の分析ではさらに、発見能力の不足による失敗は特に修復しやすいことが示される。 最後に、これらの知見に基づき、基本要素によって導かれた推論を選択的に生徒モデルへ移す、基本要素を特権情報として用いる自己蒸留フレームワーク[原文表記:\abs{}]を導入する。広範な実験により、この手法は、複数のモデル規模と難しいベンチマークにわたって、ベースラインより一貫して数学推論を改善することを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions. In this paper, we take a first step toward systematically studying mathematical understanding in LLMs, from diagnosing its distinct capabilities to leveraging these findings to improve post-training. First, we introduce the notion of Mathematical Primitive to probe structural mathematical understanding and propose \hlei{}, a novel benchmark that evaluates mathematical reasoning along four distinct dimensions: Discovery, Generation, Digestion, and Execution. Second, our systematic diagnosis shows that solution accuracy masks distinct capability profiles, primitives unlock substantial latent execution capacity, and Discovery is the dominant bottleneck in mathematical reasoning. Our post-training analysis further shows that discovery-limited failures are particularly amenable to repair. Finally, building on these findings, we introduce \abs{}, a primitive-privileged self-distillation framework that selectively transfers primitive-guided reasoning into the student model. Extensive experiments demonstrate that \abs{} consistently improves mathematical reasoning over baselines across model scales and challenging benchmarks.
著者のコメント
27 pages
arXiv ID: 2610.02191 / 要約の誤りについて