arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

Webエージェントの成否と操作履歴を人が再確認する

Auditing Web Agent Evaluation on WebArena-Lite: Human Review of Outcomes and Trajectories

Chengguang Gan, Zimeng He, Yoshihiro Tsujii, Ken-ichiro Kobayashi, Hiroki Itoh, Kotaro Funakoshi

この論文をやさしく読む

ひとことで言うと

Web操作エージェントの自動採点を人が点検し、成功の見落としや、途中まで進んだのに失敗する原因を調べています。

何に役立つ?

Webエージェントを比較する際に、最終スコアだけでなく評価器の誤判定や操作途中の詰まり方を確認するために役立ちます。

この研究の面白いところ

採点の訂正と、エージェント自体への状態管理・手順案内の効果を分けて調べています。序盤の進捗が大きくても完了に結び付かない点を履歴から確認しています。

どこまで分かった?

対象はWebArena Liteの165タスクと6条件です。人による訂正後の成功率と、自動評価器のスコアは区別が必要です。Qwenモデルの「未訓練」が具体的にどの追加訓練を指すかは要旨に説明されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Webエージェントは大規模言語モデルの重要な応用だが、その評価は、最終結果だけを見る規則ベースまたは言語モデルによる評価器に依存することが多い。タスク完了の人による確認や、失敗した操作履歴の詳しい分析はまだ限られている。本研究では、GPT 5.5と未訓練のQwen3.5 9Bモデルから構成した6つの評価条件で、WebArena Liteの全165タスクを点検する。点検では元のスコアを保持し、自動評価器の偽陰性を訂正し、結果に影響した最初の誤りを特定して、操作履歴全体の進捗を調べる。 また、実行状態を明示的に保持するMemory and Analysis Support Mechanism(MASM)と、タスクに関係する操作手順の案内を提供するGuide Textを検討する。GPT 5.5の4設定では、人による確認により、評価器が見落としていた成功を5.45〜8.49パーセントポイント分回復した。25ステップの予算では、Guide Textにより、MASMを使う場合の訂正後の成功率が34.55%から38.18%に上がる。未訓練のQwen3.5 9Bモデルでは、MASMによって評価器のスコアが13.90%から18.80%に上がる。 失敗したGPT 5.5の102件の操作履歴を確認すると、スクロールのループ、未完了の探索、早すぎる回答、無効な操作、未完了のフォーム処理が多く見られた。ステップ単位の証拠からは、序盤で大きく進捗していても、最終的には失敗することがあると分かる。これらの結果は、最終スコアだけではWebエージェントの振る舞いを十分に説明できない理由を示し、人の確認を根拠とし、操作履歴を考慮した検証の必要性を示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Web agents are an important application of large language models, yet their evaluation often depends on rule based or language model evaluators that inspect only the final outcome. Human verification of task completion and detailed analysis of failed trajectories remain limited. We audit all 165 WebArena Lite tasks under six evaluation conditions built from GPT 5.5 and an untrained Qwen3.5 9B model. The audit retains the original score, corrects false negatives from the automatic evaluator, identifies the first consequential error, and examines progress across the trajectory. We also study a Memory and Analysis Support Mechanism (MASM), which maintains explicit execution state, and Guide Text, which provides task relevant procedural guidance. Across four GPT 5.5 settings, human review recovers 5.45 to 8.49 percentage points of success missed by the evaluator. With a 25 step budget, Guide Text raises corrected success with MASM from 34.55% to 38.18%. On the untrained Qwen3.5 9B model, MASM raises the evaluator score from 13.90% to 18.80%. Review of 102 failed GPT 5.5 trajectories reveals frequent scrolling loops, unfinished exploration, premature answers, invalid actions, and incomplete form workflows. Step level evidence further shows that substantial early progress can coexist with a final failure. These results show why final scores alone provide an incomplete account of web agent behavior and motivate human grounded, trajectory aware verification.

著者のコメント

13 pages, 1 figure, 10 tables. Accepted as a poster at the NeurIPS 2026 Workshop "Who Verifies the Agents? Toward Reliable Agent Development"

arXiv ID: 2610.01491 / 要約の誤りについて