利用者も環境を変更できる条件でAIの安全性を測る
DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
この論文をやさしく読む
ひとことで言うと
AIだけでなく利用者も環境を変更できる状況を作り、やり取りの途中で安全性がどう変わるかを測るテストです。
何に役立つ?
モデル単体の応答テストでは見落とす、ツールや環境との相互作用に伴う安全性の問題を評価するのに役立ちます。
この研究の面白いところ
利用者側にも環境を動かす能力を与えたところ、実験全体の攻撃成功率が26.9%から41.1%へ変わりました。環境の制御条件そのものが評価を左右することを示します。
どこまで分かった?
数値は指定したベンチマーク、14モデル、8分野での結果です。個々の製品の実運用における被害率や、すべての攻撃に対する成功率を表すものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
LLMに基づくエージェントは、ユーザー、ツール、外部システムと相互作用する環境で動作することが増えている。しかし、安全性評価の多くは受動的なユーザーと固定された制御を仮定し、実際のエージェントの振る舞いを形づくる対話的な変化を無視している。 エージェントとユーザーの両方が共有環境の状態へ影響できる「二重制御」の対話のもとで、エージェントの安全性を測るベンチマークと評価手順、DUMA-Benchを導入する。DUMA-Benchはτ²-benchを拡張し、RAG汚染、エージェント間の操作、安全でない出力の取り扱いなど、8種類の脆弱性クラスを含む敵対的環境を備える。 OpenAI、Anthropic、DeepSeek、Qwen、Z.aiという5系統の14モデルを、8分野と複数のユーザー行動条件で評価する。実験全体で、二重制御の対話を導入すると、攻撃成功率は26.9%から41.1%へ上昇した。これらの結果は、エージェントの安全性がモデルだけの性質ではなく、モデル、ユーザー、環境の相互作用から生じることを示す。DUMA-Benchは、現実的なエージェント配備の安全性を研究するために不足していた評価の層を提供する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignoring the interactive dynamics that shape real agent behavior. We introduce \textbf{DUMA-Bench}, a benchmark and evaluation protocol for measuring agent security under \emph{dual-control} interaction, where both the agent and the user can influence the shared environment state. DUMA-Bench extends $\tau^2$-bench ~\cite{barres2025tau} with adversarial environments covering eight vulnerability classes, including RAG poisoning, cross-agent manipulation, and unsafe output handling. We evaluate \textbf{14 models from five model families} (OpenAI, Anthropic, DeepSeek, Qwen, and Z.ai) across eight domains and multiple user-behavior regimes. Across our experiments, introducing dual-control interaction increases the attack success rate from \textbf{26.9\%} to \textbf{41.1\%}. These results show that agent security is not solely a property of the model but emerges from the interaction between the model, the user, and the environment. DUMA-Bench provides a missing evaluation layer for studying security in realistic agent deployments.
著者のコメント
ACL ARR 2026 March Findings
arXiv ID: 2609.24662 / 要約の誤りについて