arXiv論文メモ
新着一覧
cs.SE / cs.AI / cs.LG · 査読状況未確認

コーディングエージェントの低コストA/B評価

DeltaSelect: Affordable A/B Testing for Coding Agents

Nicholas J. Conn

この論文をやさしく読む

ひとことで言うと

コード作成エージェントの改良を頻繁に比較できるよう、全体成績を追いやすい小さな課題集合を予算内で選ぶ方法です。

何に役立つ?

開発中の指示やスキルのA/B比較に使うための評価で、モデルの総合ランキングを作る目的ではありません。

この研究の面白いところ

公開試行の再標本化から、安定して全体を追える課題は113件中22件と分かりました。課題選択と成績較正を組み合わせ、事例では13評価を計27.86ドルで実施しています。

どこまで分かった?

採用版の費用は58.1%下がり有意でしたが、較正スコア上昇のp値は0.326です。費用改善と精度向上の統計的な根拠を同じ強さで扱うことはできません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

コーディングエージェントのベンチマークは包括比較向けで、開発中の頻繁な意思決定には向かない。DeepSWEの公開試行の再サンプリングでは、全ベンチマーク性能とのPearson相関の5パーセンタイルが0.50以上の課題は113件中22件、19.5%だけだった。本論文は、1回の結果が全体性能を追跡する課題をPearson相関で選び、分数検証結果を線形回帰で共通スコアに写像し、予算内で固定課題集合を選ぶオープンソース手法DeltaSelectを提示する。これはモデル順位付けでなく開発時のベースライン対候補の反復比較用である。gpt-5.6-luna低推論設定の事例ではカスタムスキルと指示の改訂に使った。13回の評価で記録コストは27.86米ドル。採用版は初期版より58.1%安く、1.75対4.18米ドル(p=0.008)で、較正済みスコアは42.36対36.46%と高かったが、公開版相当分散のp値は0.326だった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance. The paper presents DeltaSelect, an open-source method that identifies tasks whose one-run results consistently track full-benchmark performance using Pearson correlation, maps fractional verifier results to a common score using linear regression, and selects a fixed task set within a dollar budget. DeltaSelect is intended for repeated baseline-versus-candidate comparisons during development, not model rankings. In a gpt-5.6-luna low-reasoning case study, DeltaSelect was used to revise custom skills and instructions. Across 13 evaluations, the recorded cost was USD 27.86 at rates published August 16, 2026. The adopted version cost 58.1% less than the initial version (USD 1.75 versus USD 4.18; p=0.008), while the calibrated score was higher (42.36% versus 36.46%; published-analog variance p=0.326).

arXiv ID: 2609.19607 / 要約の誤りについて