AI研究エージェントが実験結果を理解できるか測る
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
この論文をやさしく読む
ひとことで言うと
AIが少数の実験から、設定変更が成果に与える影響をどれだけ正確に予測できるか測る評価基盤。
何に役立つ?
実験を行うAIエージェントの能力や、次に何を測るかを決める手法の比較に使える。
この研究の面白いところ
単なる最良設定の発見だけでなく、各設定変更の効果を全組合せについて予測させる。
どこまで分かった?
報告された改善値は記載された課題、データ源、測定回数での結果であり、あらゆる研究分野への一般化は示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
AI研究エージェントには、自分の実験が結果をどう変えるかについて信頼できる知識が必要である。本研究は、限られた予算で実験した後に、構成要素の変更による結果をどれだけ正確に予測できるかという実験理解を測るWhatWorkedBenchを導入する。エージェントはコードを調べ、測定を選び、構成要素の全ての設定組合せについて得点を予測する表である応答曲面を提出する。CPUでの全構成の実行から、他の要素を固定して各要素を変えたときの参照効果を得る。これらの効果は30のデータ源と8種類の作業手順にまたがる36課題、1,248件の構成記録における変更の組合せを含む。中核となる評価は、8群全体の4,206件の数値対照記録と、元の6群における108件のエージェント実行を組み合わせる。 新たな測定を8件行った場合、要素の組合せの効果を扱うリッジ法は22のデータ源のうち15で最適設定を選び、3つでは全ての効果の誤差を得点範囲の10%以内に抑えた。同じエージェントの観測へガウス過程を当てはめると、真の効果の大きさに対する精度を表す効果回復率は、元のFlash集団で0.632から0.698に、追加集団で0.621から0.720に上がった。完了した拍動検出とグラフの提出6件では、同じ観測を用いたガウス過程が群ごとのマクロ回復率を0.303から0.455に上げた。六つの二値選択肢を持つ六つの作業手順で新たな測定を20件行う場合、同じ振る舞いをする構成をコードから同値として符号化すると、ガウス過程の回復率は0.248から0.462に上がった。このベンチマークは、実験を行うエージェント、適応的な実験設計、数値推論、プログラム構造の利用に関する研究を支援する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions about component changes after budgeted experimentation. Agents inspect code, select measurements, and submit a response surface, a table predicting scores for every configuration of component settings. Exhaustive CPU execution supplies reference effects for changing each component while holding the others fixed. These effects capture combinations of changes across 36 tasks from 30 data sources and 8 workflow types, with 1248 configuration records. Core evaluation combines 4,206 numerical-control records across all eight families and 108 agent episodes across the original six. At eight new measurements, pair-effect ridge selects an optimum on 15 of 22 sources and limits every effect error to 10% of score range on three. Fitting a Gaussian process (GP) to the same agent observations raises effect recovery, accuracy relative to true effect magnitude, from 0.632 to 0.698 in the original Flash cohort and from 0.621 to 0.720 in an additional cohort. On six completed beat-detection and graph submissions, the same-observation GP raises family-macro recovery from 0.303 to 0.455. On six workflows with six binary options at 20 new measurements, encoding code equivalences, configurations with identical behavior, raises GP recovery from 0.248 to 0.462. WhatWorkedBench supports research on experimental agents, adaptive experimental design, numerical inference, and use of program structure.
arXiv ID: 2609.27490 / 要約の誤りについて