実行可能な仕様から自動プログラム修復の評価問題を生成
Specification-Driven Benchmarking for Automated Program Repair From Static Corpora to Executable Specifications
この論文をやさしく読む
ひとことで言うと
自動プログラム修復の評価問題を、固定の問題集ではなく実行可能な条件から生成する設計を提案した。
何に役立つ?
修復手法の評価で難易度や欠陥の種類を調整し、問題を再生成する仕組みを設計する参考になる。
この研究の面白いところ
問題の生成と検証を独立させ、仕様で宣言した性質を別途確かめられる構成を重視する。
どこまで分かった?
要旨が示すのは概念的な枠組みと一連の例であり、既存ベンチマークに対する大規模な比較結果は記されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自動プログラム修復(APR)のベンチマークは従来、含まれる欠陥の特性を引き継いだ固定のデータ集合として作られてきた。この方法は数十年にわたる進歩を支えたが、有限の問題集では実験条件の制御が限られ、繰り返し使うと評価対象への混入が起こりやすくなり、評価要件の変化に合わせた体系的な再生成や調整もできない。本研究は、実行可能な仕様でベンチマークを定義し、問題生成を通じて具体化する「仕様駆動型ベンチマーキング」を提案する。仕様では、プログラムの文脈、欠陥の分類、難易度、検証方法、問題集の制約など、求める性質を明示する。生成パイプラインは、独立した生成、検証、問題集管理の要素を使って、それらの要件を実現する。 本研究はこの方法の概念的な基礎として、ベンチマーク仕様の各側面の分類体系を導入し、各側面が決定的な構成上の責任へどう対応するかを定め、信頼できる問題生成には独立した検証が構造上必要だと論じる。一連の作業例を通じて、仕様上の選択がパイプラインに伝わり、性質を独立に検証できる評価問題が作られる様子を示す。固定のデータ集合ではなく実行可能な仕様としてベンチマークを扱うことで、評価問題の構築を、既存の成果物を集める作業から、条件を宣言する実験設計へ移す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Automated Program Repair (APR) benchmarks have traditionally been constructed as static datasets whose characteristics are inherited from the defects they contain. While this paradigm has enabled decades of progress, finite corpora provide limited experimental control, become increasingly susceptible to contamination as they are reused, and cannot be systematically regenerated or adapted as evaluation requirements evolve. We propose specification-driven benchmarking, a paradigm in which benchmarks are defined by executable specifications and realized through benchmark generation. The specification explicitly declares the intended properties of the benchmark (including program context, fault taxonomy, difficulty, validation strategy, and corpus constraints) while a generation pipeline realizes those requirements through independent generation, validation, and corpus management components. We develop the conceptual foundations of this approach by introducing a taxonomy of benchmark specification dimensions, establishing how each specification dimension maps to deterministic architectural responsibilities, and arguing that independent validation is a structural requirement for trustworthy benchmark generation. An end-to-end example illustrates how specification choices propagate through the pipeline to produce benchmark instances whose properties are independently verifiable. By treating the benchmark as an executable specification rather than a static dataset, the proposed paradigm shifts benchmark construction from artifact curation to declarative experimental design.
arXiv ID: 2609.28896 / 要約の誤りについて