arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

AIエージェントの評価では情報・時間・検証手段も測る

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata

この論文をやさしく読む

ひとことで言うと

AIエージェントの成績が、モデルだけでなく、渡す情報、制限時間、検証ツールなどの設定でどう変わるかを調べています。

何に役立つ?

エージェントの評価や運用を設計するとき、モデル名と成功率だけで比較せず、設定条件と繰り返しのばらつきを記録するために役立ちます。

この研究の面白いところ

評価した条件では、情報の与え方が時間やモデル規模より大きく影響しました。検証を言葉で頼む場合と、専用ツールを渡す場合で行動が大きく違う点も示しています。

どこまで分かった?

四つの科学課題で専門モデルを探して操作するベンチマークの結果です。約54%という分散の割合や各設定の効果を、すべてのエージェント課題にそのまま一般化はできません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

エージェントの評価では、単一の成功率を超えて、コスト、一貫性、頑健性などの指標を報告することが増えている。しかし通常、エージェント自体は固定されたものとして扱われる。実際には、エージェントは設定可能なシステムである。ユーザーは何を伝えるか、どれだけ長く実行させるか、どのモデルを使うかを決め、それぞれの選択が性能とその安定性を変えうる。 本研究では、コーディングエージェントが公開された専門モデルを見つけ、正しく操作しなければならない四つの科学課題からなる新しいベンチマークで、これらの選択を調べる。エージェント構成の五つの要素、すなわち課題情報、推論、自己検証、時間予算、バックボーンモデルを検討する。実行間の大きなばらつきが見られ、結果の分散の約54%は構成を変更したことではなく、同じ構成で繰り返し実行したことに由来していた。 構成間の比較では、エージェントへ与える情報が最も大きな効果を持ち、時間予算とモデル規模の効果を上回るとともに、コストを減らし、較正も改善した。設定項目どうしにも相互作用があり、時間の追加が役立つのは、それを活用できる十分な情報または十分な能力のモデルをエージェントが持つ場合に限られた。最後に、実行軌跡に基づく行動分類から、回答を検証するよう指示しても検証行動にはほとんど影響しない一方、専用の検証ツールを与えると行動が大きく変わることが分かった。 これらの結果は、エージェント自体を設定可能なシステムとして評価すべきこと、また、望ましい行動の一部はプロンプトで要求するよりシステムへ実装した方が効果的であることを示唆する。ベンチマークと18,000件を超えるエージェントの実行軌跡を公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.

arXiv ID: 2610.01618 / 要約の誤りについて