arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

結果と行動過程からエージェント評価を小さくする

Beyond Outcomes: Dual-View Relational Learning for Efficient Agent Benchmarking

Xinshuai Guo, Junjie Wu, Dolly Deng, Yinghui Li, Hai-Tao Zheng, Suncong Zheng, Maxm Pan

この論文をやさしく読む

ひとことで言うと

少数の課題でエージェント全体の評価を推定するため、最終点数だけでなく途中の行動過程も使って課題を選びます。

何に役立つ?

評価コストを抑えながらモデルの得点や順位を比較し、能力の違いを確認する用途があります。

この研究の面白いところ

6種類の過程信号を使い、結果と過程の2つの関係からちょうど必要数の課題を選ぶ仕組みです。20課題で24〜40倍の圧縮を報告しています。

どこまで分かった?

性能は5ベンチマーク上の比較結果です。MAEや順位相関の改善であり、選んだ20課題があらゆる新しいエージェントを誤差なく評価できる保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

エージェントのベンチマーク評価は、通常のLLMベンチマークより大幅に高コストである。そのためベンチマークの圧縮は自然な解決策となるが、既存手法は主にタスクとモデルの最終得点分布にある冗長性をモデル化しており、これはエージェント評価でも重要である。この限界に対応するため、大規模な行動軌跡を解析し、最終的なエージェント性能と系統的に関連する6つの相補的な過程信号を特定する。より完全な観点から性能の冗長性を解きほぐすため、結果と過程の関係を共同で使い、指定された件数ちょうどの小さな評価集合を学習し、全ベンチマークの得点を予測する圧縮手法DualViewEvalを提案する。 5つのエージェントベンチマークと5つの代表的な比較手法に対し、DualViewEvalはすべてのデータセットで最良の結果を達成する。わずか20タスクでAPEX-AgentsとBFCLを24〜40倍に圧縮し、最も強い比較手法より平均絶対誤差(MAE)を14.5〜28.2%削減する。また、SWE-bench VerifiedではEssenceBenchに対し、Kendallのτを最大7.2%改善する。選ばれた小規模集合は、エージェント間の能力差も明らかにし、効率的なエージェントモデル開発へ、簡潔で診断的なフィードバックを提供する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Agent benchmarks are substantially more costly to evaluate than conventional LLM benchmarks. Benchmark compression is therefore a natural solution, yet existing methods primarily model redundancy in task--model final-score distributions, which is important in agentic evaluation. To address this limitation, we analyze large-scale trajectories and identify six complementary process signals that are systematically associated with final agent performance. To disentangle agent performance redundancy from a complete perspective, we propose DualViewEval, an agent benchmark compression method that jointly exploits outcome and process relations to learn an exact-size miniset and predict the full-benchmark scores. Across five agent benchmarks and five representative baselines, DualViewEval achieves the best results in all datasets. With only 20 tasks, it achieves $24\times$--$40\times$ compression on APEX-Agents and BFCL, reducing mean absolute error (MAE) by $14.5\%$--$28.2\%$ over the strongest competitors while improving Kendall's $\tau$ by up to $7.2\%$ relative to EssenceBench on SWE-bench Verified. The selected minisets further reveal capability differences among different agents, providing compact and diagnostic feedback for efficient agentic model development.

arXiv ID: 2609.18909 / 要約の誤りについて