証明に失敗したとき数学概念からLean用スキルを更新
SkillEvoLean: Mutation-enhanced skill evolution for Lean provers
この論文をやさしく読む
ひとことで言うと
Leanで証明を作るAIが行き詰まったとき、数学の概念を手掛かりに解き方や参照知識を作り替える方法です。モデル自体の重みは更新しません。
何に役立つ?
証明の候補がすべて失敗し、成功例から改善できない場面で、次に試す方針を作るために役立つ可能性があります。実証対象は記載された定理証明ベンチマークと数学競技問題です。
この研究の面白いところ
指示文だけでなく、参照する概念や技法も更新します。ランダムな文章による変異と比べることで、数学概念を使うこと自体の効果も調べています。
どこまで分かった?
成功率はGPT-5.5と指定された評価条件での結果です。IMOとUSAMOはいずれも6問中4問であり、全問証明ではありません。他モデルや任意の形式化された定理への一般保証は要旨にありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
スキルの進化は、大規模言語モデルのエージェントをパラメータ更新なしで改善する有望な方法だが、形式的な定理証明への利用は十分に研究されていない。既存の方法は主に自然言語での推論を対象とし、成功・失敗の実行軌跡を解析し、解法戦略を段階的に修正することでスキルを改善する。Leanの検証器は信頼できる実行フィードバックを与えるが、サンプリングした実行軌跡がすべて失敗した場合、既存のスキル進化法には、有効な更新方向を推測するための成功軌跡がない。さらに、これらの方法は主に最上位の指示ファイルに注目し、数学的概念や証明技法を含む参照知識の進化を十分に探索していない。 これらの限界に対処するため、スキルで強化されたLean証明器を構築する、変異を取り入れたスキル自己進化の枠組みを提案する。この枠組みは、漸進的な更新と変異に基づく更新を通じ、高水準の解法方針と参照知識を共に進化させる。漸進的な進化では成功・失敗の実行軌跡から局所的な改善を導く。一方、完全な証明を生成できない場合には変異が起動し、数学的概念をサンプリングして新たなスキル候補を作り、検証器のフィードバックの下で選択する。 MiniF2F、PutnamBench、2025年国際数学オリンピック(IMO 2025)、2026年米国数学オリンピック(USAMO 2026)で本手法を評価する。基盤モデル、実行軌跡のサンプリング予算、テスト時の計算量を同じにした条件で、GPT-5.5を用いる本手法はそれぞれ100.0%、90.6%、4/6、4/6の証明成功率を達成し、ベースライン手法を上回る。追加解析では、概念に導かれた変異は、ランダムな文章に導かれた変異より、MiniF2Fで6.9ポイント、PutnamBenchで8.2ポイント高く、IMO 2025とUSAMO 2026でもそれぞれ1問多く解いた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Skill evolution offers a promising way to improve large language model agents without updating their parameters, but its use in formal theorem proving remains underexplored. Existing methods mainly target natural-language reasoning, improving skills by analyzing successful and failed trajectories and incrementally revising solving strategies. Although the Lean verifier provides reliable execution feedback, when all sampled trajectories fail, existing skill evolution methods lack successful trajectories from which to infer effective update directions. Furthermore, these methods also focus mainly on the root instruction file, thus underexploring the evolution of reference knowledge including mathematical concepts and proving techniques. To address these limitations, we propose a mutation-enhanced skill self-evolution framework for building skill-augmented Lean provers. The framework jointly evolves a high-level solving policy and its reference knowledge through progressive and mutation-based updates. Progressive evolution derives local improvements from successful and failed trajectories, while mutation is triggered when no complete proof can be generated, sampling mathematical concepts to produce and select new skill candidates under verifier feedback. We evaluate our method on MiniF2F, PutnamBench, the 2025 International Mathematical Olympiad (IMO 2025), and the 2026 USA Mathematical Olympiad (USAMO 2026). Under the same backbone model, trajectorysampling budget, and test-time compute, our method achieves proof success rates of 100.0%, 90.6%, 4/6, and 4/6, respectively, with GPT-5.5, outperforming the baseline methods. Further analysis shows that concept-guided mutation outperforms random-text-guided mutation by 6.9 and 8.2 percentage points on MiniF2F and PutnamBench, respectively, while solving one additional problem on both IMO 2025 and USAMO 2026.
arXiv ID: 2610.01799 / 要約の誤りについて