全モデル失敗のベンチマーク課題は本当に難しいか
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
この論文をやさしく読む
ひとことで言うと
全てのAIエージェントが失敗した課題を、実際に解けない課題として扱えるか、運用記録を調べた。
何に役立つ?
ベンチマークを作成・利用する際、成功率ゼロの原因を切り分けるのに役立つ。実際の記録では125課題中78課題だけが未解決候補として認定された。
この研究の面白いところ
模範解答、空の解答、設備の記録、回避試行などを組み合わせて失敗の原因を審査した点。単一の成功率から能力の限界を読み取る危うさを数で示した。
どこまで分かった?
「認定された未解決」は評価したエージェントの失敗と一定の妥当性条件を示すだけで、課題の本質的な難しさや検証器の完全性を証明しない。分析対象は固定された運用記録である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
最先端モデルを評価するには、現在のモデルが解けない課題が必要である。ただし、どのモデルも解けない課題が必ずしも本質的に難しいとは限らない。成功率ゼロは能力不足だけでなく、文脈の欠落、壊れた模範解答、基盤設備の障害、回避できる検証器でも起こる。本研究は、Terminal-Bench 3/Frontier-Bench 0.1の固定された運用記録を分析した。記録にはプルリクエスト1,081件、採点対象639課題、試行28,801回、記録されたエージェント費用105,933ドルが含まれる。正当な合格が一つもない125課題について、課題資料、模範解答の実行、空の解答を使う対照試験、敵対的試行、実行過程、計測記録、レビュー記録を組み合わせ、順序を定めた妥当性審査を行った。その結果、未解決の候補として認定できたのは78課題だけだった。残りには、判定基準の破綻14件、基盤設備の障害が支配的な8件、検証器の回避によってしか合格できない4件、利用可能な証拠では解けるかどうかを認定できない21件が含まれた。「認定された未解決」という表示の意味は限定的で、作成者が想定した解法は通り、設備障害は支配的でなく、明確な回避は観測されず、評価した全エージェントが失敗したということである。課題固有の難しさ、検証器の完全性、狙った能力での失敗まで証明するものではない。不合格の提出物と合格した課題も分析し、成功率だけでは難しさの理由を説明できないと示した。著者らは、ベンチマークが全失敗の課題を能力に関する主張へ使う前に、その根拠を報告すべきだと結論づける。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be bypassed. In this paper, we study this issue using a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record with 1,081 pull requests, 639 scored tasks, 28,801 trials, and $105,933 in logged agent spend. We ask what an all-fail task actually certifies. For the 125 tasks with no honest pass, we combine task artifacts, reference-solution runs, empty-solution controls, adversarial trials, trajectories, telemetry, and review records, and apply an ordered validity screen. Only 78 of the 125 tasks survive as certified-unsolved candidates. The remaining tasks include 14 with broken oracles, 8 dominated by infrastructure failures, 4 that are only passable through verifier bypasses, and 21 whose solvability is not certified by the available evidence. Thus, lack of saturation and genuine difficulty are not the same thing. The certified-unsolved label is also narrow: it means that the authored route passed, infrastructure did not dominate, no strict bypass was observed, and all evaluated agents failed. It does not prove intrinsic hardness, verifier completeness, or failure at the intended capability. We further analyze rejected submissions and passing tasks to show that pass rate alone cannot explain why a task is difficult. Overall, our results suggest that frontier benchmarks should report the evidence behind their all-fail tasks before using them as capability claims.
arXiv ID: 2609.26826 / 要約の誤りについて