arXiv論文メモ
新着一覧
cs.AI / cs.IT / math.IT · 査読状況未確認

限られた評価回数でベンチマークの不確かさを測る条件

Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yilan Wei, Yankai Zeng, Bojun Lin

この論文をやさしく読む

ひとことで言うと

AIベンチマークを限られた回数で評価するとき、課題を広く試すか同じ課題を繰り返すかを調べた研究です。

何に役立つ?

評価予算を配分し、得点だけでなく信頼区間の幅も管理する際の参考になります。

この研究の面白いところ

理論上の最適な区間幅を示し、LiveCodeBenchの再評価では点推定誤差と信頼区間の幅をともに減らしています。

どこまで分かった?

実証結果は16モデル、880課題、各課題五出力のLiveCodeBench再評価でのものです。理論上の率には要旨に示されたLとαの条件があります。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

評価を繰り返せばベンチマークの得点を正確に推定できても、不確かさが小さいと保証するには反復がさらに必要になる場合がある。本研究は、M個の課題それぞれにL個の二値の経路がある固定した配置で、その必要条件を特徴付ける。総予算は(M+t)Kで、各経路に必要な応答またはエピソードは最大K個とする。L≥3、0<α≤1/12を固定した場合、最悪の純粋な群での信頼区間の最適な期待幅は、すべての課題を観測する場合には[M(t+1)]^{-1/2}のオーダー、課題の省略を許す場合には[M(t+√M)]^{-1/2}のオーダーとなる。ただし定数はαとLに依存する。下限は固定予算内で適応的に選ぶ方策も対象とし、無作為に固定した部分集合を使う設計は、不一致による証明を通じて両方の率を達成する。平均値と不一致を合わせた区間により、課題の網羅に関する法則を、有限予算で使える推論に変える。16モデル、880課題、各課題五つの出力による同一予算のLiveCodeBench再評価では、課題を網羅する設計により、まとめて一様抽出する方法に比べて点推定の中央値の平均二乗誤差が87.0%下がった。また、Joint証明による信頼区間は16区分中15区分で狭く、中央値の区間幅が30.6%下がった。有限の条件での解析により、評価した規模では課題を広く対象にすることが有効な選択だと分かり、群の大きさと課題内の一致度が有用な運用範囲をどう決めるかも特徴付けた。理論上の鋭い法則と固定予算での証拠を合わせ、反復数と課題の網羅を、情報効率の良い繰り返し評価の明示的な設計要素とする。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of $M$ tasks with $L$ binary paths per task under the hard budget $(M+t)K$, where each path costs at most $K$ responses or episodes. For fixed $L \ge 3$ and $0 < \alpha \le 1/12$, the optimal expected width on the worst pure cohort is $\Theta_{\alpha,L}([M(t+1)]^{-1/2})$ when every task is observed and $\Theta_{\alpha,L}([M(t+\sqrt{M})]^{-1/2})$ when omission is allowed. The lower bounds cover adaptive hard-budget policies, and fixed random-subset designs attain both rates through disagreement certificates. A joint mean/disagreement interval turns the task-covering law into practical finite-budget inference. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task-covering design reduces median point-estimation MSE by 87.0\% relative to pooled uniform sampling, while the Joint certificate produces narrower confidence intervals in 15/16 panels and reduces median interval width by 30.6\%. Finite-regime analyses identify task coverage as the effective choice at the evaluated scale and characterize how cohort size and within-task agreement determine the useful operating region. Together, the sharp laws and fixed-budget evidence make replication and task coverage explicit design variables for information-efficient repeated evaluation.

arXiv ID: 2609.29140 / 要約の誤りについて