arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

誤った数学命題への迎合を減らすモデルEuston

Euston: Training Away Mathematical Sycophancy Without Losing the Mathematics

Zehua Cheng, Wei Dai, Jiahao Sun

この論文をやさしく読む

ひとことで言うと

誤った数学命題を証明するよう求められたとき、無理に証明を作らず判定するモデルを学習する。

何に役立つ?

数学の推論モデルが誤った前提に迎合する問題を評価・改善する方法の参考になる。

この研究の面白いところ

対応する真偽の命題対を使い、識別力を上げつつAIME成績には統計的に有意な低下がなかった。

どこまで分かった?

公式評価セットがすべて偽であることや、実際の誤り出現率での適合率の低さが、結果の解釈を制限する。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

推論を行う言語モデルは解答を作るよう訓練されており、誤った問題を渡されたときにもその傾向が残る。強いモデルに改変された定理の証明を求めると、通常は応じて、誤った内容を自信ありげに導く。本論文は、まさにこれに抵抗するよう学習した、80億パラメータの数学的主張検証モデルEustonを示す。訓練データは、属性の多様性と生成時の構造マスク、範囲を同期させた検証を組み合わせる確率的な因子グラフ生成器GraphSynthで作った。2010~2025年のarXiv論文から、真の主張と改変した主張の対応する3026組、計6052文を得た。DeepSeek-R1-8Bを、規則に基づく外部APIを使わない報酬でGRPOによって、H100 GPU四基を使い189段階追加学習した。 真200件・偽200件の均衡した未使用データでは、均衡正解率は29.50%から63.75%へ上がった。偽を偽と判定する割合から真を偽と判定する割合を引いた識別の差は、−0.5%(z=−0.1)から+27.5%(z=+6.0)へ改善した。この改善は一般的な数学能力を大きく犠牲にしていない。公式の判定条件によるAIME 2026の正答率は、元の69.17%に対し65.00%で、差の−4.17%は統計的に有意ではなかった。一方、より小さなGraphSynthデータを使った同じ手法の以前の実行では40.00%まで低下していた。 回答の長さの中央値は19217トークンから18296トークンへ、途中切断率は25.8%から8.3%へ下がり、長く考えたための改善ではない。著者らは、結果の解釈を制限する交絡要因も報告する。主なものは、公式評価セットがすべて偽の命題からなることと、現実の誤りの出現率では適合率が低くなる可能性である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Reasoning language models are trained to produce solutions, not to refuse them, and this bias persists when the problem they are handed is false. Asked to prove a corrupted theorem, a strong model will typically comply and produce a confident derivation of something untrue. We present Euston, an 8B mathematical claim-verification model trained to resist exactly this. Training data were generated with GraphSynth, a probabilistic factor-graph generator that couples attribute-level diversity to decode-time structural masking and span-synchronized verification, yielding 3{,}026 matched true/corrupted statement pairs (6,052 statements) drawn from arXiv papers spanning 2010--2025. We fine-tuned DeepSeek-R1-8B with GRPO under a rule-based, zero-API reward for 189 steps on four H100 GPUs. On a balanced 200-true/200-false held-out split, balanced accuracy rises from 29.50% to 63.75% and the discrimination gap---the difference between the rate of calling false statements false and the rate of calling true statements false moves from -0.5% (z=-0.1) to +27.5% (z=+6.0). Critically, the gain is not purchased with general mathematical ability: AIME 2026 accuracy under official semantics is 65.00% against a 69.17% base, a difference of -4.17% that is not statistically significant, whereas an earlier run of the same recipe on a smaller GraphSynth corpus collapsed to 40.00%. Median response length also falls from 19,217 to 18,296 tokens and the truncation rate from 25.8% to 8.3%, so the improvement does not come from thinking longer. We report the result together with the confounds that bound its interpretation, principally the all-false composition of the official evaluation sets and the low precision implied at realistic error prevalence.

arXiv ID: 2609.23205 / 要約の誤りについて