FORMのコードを実行検証で学習させる
LLM-Based FORM Code Generation with Verification-Driven Fine-Tuning
この論文をやさしく読む
ひとことで言うと
粒子物理の記号計算言語FORMを、実際に動かして正しさを確認した例で小型モデルへ教えた。
何に役立つ?
FORMのコード作成を支援し、専門言語向けモデルの検証付き訓練を設計するのに役立つ。
この研究の面白いところ
FORM本体を正解確認に使って4,633件を作り、一回試す評価で大きな汎用モデルを上回った。
どこまで分かった?
成績は研究の四つのベンチマークと単発実行の条件による。小さく難しい課題では先端モデルとの差は統計的に確認できなかった。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
FORMは、多重ループのFeynman図の計算で生じる非常に大きな代数式を扱う、粒子物理で広く使われる記号計算言語である。しかし、著者らの知る限りFORMコードの作成を助けるAIの道具はなかった。現代の大規模言語モデルは、数千億パラメータの先端モデルを含め、文書なしで一回だけ試す条件では、指示に従う課題と入門用のFORM課題で実行成功率が0%だった。研究時点ではFORMが言語モデルにとって真のゼロショット言語であることを示す。そこでFORMの実行ファイル自体を正解確認に使い、決定的な計算、自由度のあるプログラム、入門用コード、知識の質問と回答を含む4,633件の訓練例を生成・検証する手順を提案する。小さな公開重みモデルQwen3-8Bを、量子化した低ランク適応QLoRAで微調整した。四つの補完的なベンチマーク、合計840課題を一回ずつ試す評価では、最大756Bパラメータの先端モデルを、大きなベンチマークでの実行率とFORMによる厳密な出力一致率で明確に上回った。小さく難しいベンチマークでは、統計的に区別できない成績だった。一般的な推論とコーディングの能力は、2.6パーセントポイント以内の差で保たれた。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising from multi-loop Feynman diagram calculations. Despite its central role in precision theoretical physics, no artificial-intelligence tooling exists, to our knowledge, for assisting physicists in writing FORM code. We show that contemporary large language models (LLMs), including frontier models with hundreds of billions of parameters, achieve a zero-percent execution pass rate on our instruction-following and tutorial-style FORM tasks without documentation in a single attempt, establishing FORM as a genuine zero-shot language for LLMs at the time of writing. We then present a verification-driven data generation pipeline that uses the FORM binary itself as an execution oracle to produce and validate a corpus of 4,633 training examples spanning deterministic computations, open-ended programs, tutorial code, and knowledge question-answer pairs. Fine-tuning a compact open-weights model (Qwen3-8B) with quantized low-rank adaptation (QLoRA) yields a specialist that, evaluated on four complementary benchmarks (840 tasks, single attempt each), decisively outperforms frontier models with up to 756B parameters in execution rate and in strict, FORM-verified output matching on the larger benchmarks, and remains statistically indistinguishable from them on the smaller, harder ones. General reasoning and coding capabilities are preserved within 2.6 percentage points.
arXiv ID: 2609.23367 / 要約の誤りについて