販売サイトの誘導下でも購入目的を守るAIの評価
CAVEAT: Towards Robust Computer-Use Agents in Incentive-Misaligned Environments
この論文をやさしく読む
ひとことで言うと
販売サイトが商品へ誘導する状況で、購入エージェントが利用者の希望を守れるか調べた。
何に役立つ?
利用者の代わりに購入するエージェントを評価し、環境による誘導に対処する設計に役立つ。
この研究の面白いところ
誘導を有効にすると最適商品の購入率が78.6%から17.3%へ下がり、失敗の入り口を三つに分けて分析した。
どこまで分かった?
結果は九つの市場環境と五つのモデル系列を使った評価であり、実際の購入サイト全般での割合ではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
コンピューター操作エージェントは、オンラインで利用者の代わりに行動することが増えている。では、操作先の環境の利害が利用者の利害と一致しないとどうなるか。例えばオンライン市場では、プラットフォームが特定の商品を優遇し、エージェントを利用者の目的から逸らす可能性がある。既存の評価基準は協力的な設定や明示的な攻撃を扱うが、環境自体に結果への利害があるときに利用者の目的を守れるかは試していない。本研究は、九つの市場環境と八種類の一般的な誘導手法を含む、条件を制御した評価基準CAVEATを導入する。五つのモデル系列で、条件を合わせた対照実行では78.6%の回で利用者にとって最適な商品を購入したが、誘導を有効にすると17.3%に下がった。より大きなモデルと推論量の増加は頑健性を改善したが、大きな失敗は残った。実行軌跡の分析と対象を絞った要素除去実験により、誘導が判断に入り込む三つの箇所を特定した。エージェントが利用者の優先順位をゆがめること、検討する選択肢を早く絞りすぎること、判断に必要な証拠が揃う前に決定することである。この診断に基づき、これらの失敗を直接狙うCAVEAT-Harnessを開発し、利用者にとって最適な購入を55.0%増やした。対象を絞った追加学習は、小型の公開モデルもさらに改善した。これらの結果は、利害の対立に対する頑健性が委任されたエージェント固有の課題であること、失敗の仕組み、狙いを定めた介入で大きく改善できることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Computer-use agents (CUAs) increasingly act on behalf of users online. What happens when the environments they operate in have incentives that do not align with the user's? In online marketplaces, for example, platforms may favor some products over others, potentially steering agents away from the user's objective. Existing CUA benchmarks cover cooperative settings or explicit attacks, but do not test whether agents preserve user objectives when the environment itself has a stake in the outcome. We introduce CAVEAT, a controlled benchmark spanning nine marketplace environments and a taxonomy of eight common steering mechanisms. Across five model families, agents purchase the user-optimal product in 78.6% of matched-control episodes but only 17.3% when steering mechanisms are enabled. Larger models and increased reasoning improve robustness, but substantial failures persist. Our trajectory analysis and targeted ablations identify three points where steering enters the decision process: (1) agents distort the user's priorities, (2) prematurely narrow the set of alternatives they consider, and (3) commit before resolving decision-relevant evidence. Guided by this diagnosis, we develop CAVEAT-Harness, which directly targets these failure modes and raises user-optimal purchasing by 55.0%. Targeted post-training further improves a smaller open model. These results establish incentive robustness as a distinct challenge for delegated agents, diagnose how it fails, and show that targeted interventions can substantially improve it.
arXiv ID: 2609.27273 / 要約の誤りについて