arXiv論文メモ
新着一覧
cs.CR · 査読状況未確認

AIエージェント攻撃成功率の測定方法を検証

Attack Success Rate Is Not a Number: On Measurement Validity in Agentic AI Security Evaluation

Chetan Pathade, Prathamesh Pawar, Shubham Patil

この論文をやさしく読む

ひとことで言うと

AIエージェントへの攻撃成功率が研究ごとに同じ意味で測られているかを、259本の論文と統計解析で調べた。

何に役立つ?

攻撃・防御評価の報告項目を整え、異なる論文の結果を比較できるか判断する材料になる。

この研究の面白いところ

人手分類した50本の58%が分散推定も反復実行も報告せず、100事例では5ポイントの真の差が1回の評価で約21%の確率で逆転すると解析している。

どこまで分かった?

文献調査は2025年2月から2026年9月のarXiv論文259本を対象とする。比較不能性は現在の報告状況についての結論であり、個々の研究結果を否定するものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

攻撃成功率(ASR)は、大規模言語モデルのエージェントに対する攻撃や防御を評価する論文のほぼすべてで主要な指標として使われている。しかし著者らは、現在使われるASRは単一の量ではなく、6つの設計上の選択に左右される指標の集まりであり、論文ではそれらが明記されることが少なく、研究間でも一定に保たれていないと主張する。独自システムへのアクセスを必要としない二つの研究で、この見方を裏づける。 第一に、2025年2月から2026年9月までにarXivへ投稿されたエージェントのセキュリティ論文259本を全文メタ分析した。主要な攻撃指標について、分散の推定値も反復実行も報告しない論文が多数を占めた。無作為抽出した50本を人手で分類すると58%(95%信頼区間44~71%)、全259本を自動分類すると65.3%だった。評価が確率的であったか判断できるほど復号設定を公開した論文は30.9%のみで、LLM判定器を使うと確認できた64本のうち、人間のラベルとの一致を何らかの形で確認した論文は29.7%だった。 第二に、解析的な研究により、こうした欠落が単なる形式上の問題ではないことを示す。100事例のベンチマークでは、通常の検出力で識別できるASRの最小差は18.2パーセントポイントであり、真のASRが5ポイント異なる二つの防御策を1回の評価で比べると、約21%の確率で順位を逆に判定する。また、6つの選択軸のいくつかはシステムによって異なる形でASRを変えるため、比較時に相殺される一定のずれとはならない。著者らは、現状では論文をまたいだASR比較は裏づけられていないと結論し、測定した各問題に対応する10項目の報告チェックリストを提案する。目的は個々の結果を否定することではなく、分野に欠けていた共通の測定基準を提供することにある。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Attack success rate (ASR) is the headline metric in nearly every published evaluation of attacks on, and defenses for, LLM agents. We argue that ASR as currently used is not a single quantity but a family of metrics parameterized by six design choices that papers seldom specify and never hold constant across the literature. We support this with two studies that require no proprietary access. First, a full-text meta-analysis of 259 agentic-security papers posted to arXiv between February 2025 and September 2026 finds that most report neither a variance estimate nor repeated runs for their headline attack metric: 58% (95% CI 44-71) in a hand-coded random sample of 50, 65.3% by automated coding of all 259. Only 30.9% disclose enough about decoding to establish whether their evaluation was even stochastic, and of the 64 papers we confirm use an LLM judge, 29.7% report any agreement check against human labels. Second, an analytical study shows that these omissions are not cosmetic: on a 100-instance benchmark, the minimum difference in ASR detectable at conventional power is 18.2 percentage points, and two defenses whose true ASRs differ by 5 points are ranked in the wrong order by a single-run evaluation roughly 21% of the time. Because several of the six axes shift ASR in a system-dependent way, the resulting incomparability is not a constant offset that cancels in comparison. We conclude that cross-paper ASR comparison is currently unsupported, and propose a ten-item reporting checklist targeted at each failure we measure. Our aim is not to dispute any individual result but to supply the shared measurement contract the field has so far done without.

著者のコメント

7 pages, 2 tables, 2 Figures

arXiv ID: 2609.25173 / 要約の誤りについて