arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

言語モデルの作業結果を早期予測する際の確信度を検証

Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark

YanZe Cao

この論文をやさしく読む

ひとことで言うと

言語モデルのエージェントが成功するかを途中で予測する仕組みについて、別のエージェントへ使い回したときの確信度のずれを調べた研究。

何に役立つ?

作業を早めに止めて評価費用を減らす際、予測の確信度を対象エージェントごとに確認すべきか判断する材料になる。

この研究の面白いところ

全体に共通する大きなずれは支持されない一方、特定のエージェントと予測ヘッドの組ではずれが残った。別ベンチマークでの再現性も事前登録した基準で検査している。

どこまで分かった?

持続的な誤差はSWE-bench Verifiedの固定環境内の特定の組で示された。TerminalBenchでは再現を確認できず、モデル固有またはベンチマーク一般の性質とは結論していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

エージェントの作業過程から結果を早期に予測し、十分に予測可能になった時点で実行を終了すれば、評価費用を減らせる。ただし予測器の確信度が適切に較正されていることが前提となる。学習に使っていないエージェントへ予測器を適用すると較正が崩れるおそれがあるが、その失敗が多くのエージェントに共通するのか、特定の対象エージェントと予測ヘッドの組に集中するのかは分かっていない。 公開されたSWE-bench Verifiedの作業軌跡と、固定した二つのヘッドを持つ早期結果予測処理系を使い、一つのエージェントを学習対象から外す較正監査、二つのエージェントを外す共通予測器の対照実験、真の事前確率を使う補正、学習集団・課題の再標本化・課題の前半後半・ジャックナイフ法・閾値にわたる頑健性の検査を行った。構成を固定したTerminalBenchの解析を、事前登録した適用範囲の検証とした。 同じ予測器を使った組み合わせ全体に広く異質性があるという結果は支持されなかった。補正後のギャップの組ごとの差の中央値は、成功を予測するヘッドでは0.0180(45組)、失敗を予測するヘッドでは0.0385(35組)で、どちらも事前登録した異質性の基準を満たさなかった。一方、gpt-5-miniの成功予測とclaude-opus-4.6の失敗予測という二つの特定の組では、較正を移した際の誤差が持続し、補正後のギャップの中央値はそれぞれ0.1377と0.1107だった。固定したどの対照条件でも符号は逆転しなかった。 TerminalBenchでは、異なるベンチマークで同じ結果が再現されるとは確認できなかった。成功予測の対象では判定がゼロ件で結論不能、失敗予測の対象では事前登録した持続性の基準を満たさなかった。したがって、固定された一つの環境の中では特定対象に強い較正の移転誤差が存在しうるが、その誤差がモデル固有であることや、複数ベンチマークに一般的に見られることまでは証拠が示していない。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Predicting early outcomes based on trajectory can decrease the expenses associated with agent evaluation by terminating a run once the outcome becomes sufficiently predictable, assuming that the predictor's confidence is properly calibrated. Calibration is at risk when a predictor is applied to an agent on which it was never trained, but it is not known whether such transfer failures are broad across agent systems or concentrated in specific target agent/head combinations. Using public SWE-bench Verified trajectories and a frozen dual-head early-outcome prediction pipeline, we ran a leave-one-agent-out calibration audit, a shared-predictor leave-two-agents-out control, oracle prior correction, and a robustness battery over training cohorts, task resampling, task halves, jackknife, and thresholds. Fixed-scaffold TerminalBench analysis served as a pre-registered boundary test. Broad same-predictor pairwise heterogeneity was not supported; the median pairwise corrected-gap differences were 0.0180 (SUCCESS head, 45 pairs) and 0.0385 (FAILURE head, 35 pairs), and the pre-registered heterogeneity criterion was not met on either head. Two specific combinations, gpt-5-mini/SUCCESS and claude-opus-4.6/FAILURE, showed persistent calibration-transfer errors (median corrected gaps 0.1377 and 0.1107) without a sign reversal under any frozen control. TerminalBench did not establish cross-benchmark replication: the success target produced zero decisions (INDETERMINATE), and the failure target did not satisfy the pre-registered persistence criterion. Therefore, a strong target-specific calibration-transfer error can exist within one frozen environment, but the evidence does not establish that the error is intrinsic to the model or general across benchmarks.

著者のコメント

26 pages, 4 figures, 3 tables

arXiv ID: 2609.25647 / 要約の誤りについて