AI決済の権限境界を人手の攻撃4371件で検証
APort Vault: Benchmarking AI Agent Payment Authorization with the Open Agent Passport
この論文をやさしく読む
ひとことで言うと
AIが送金を要求することと、許可されていない相手への送金が実行されることを分け、外部の権限チェックの効果を測っています。
何に役立つ?
決済エージェントの評価で、モデルの応答率だけでなく実際の権限違反を確認するために役立ちます。比較対象の範囲では、認可層を通した未許可先への送金は0件でした。
この研究の面白いところ
正常な決済を多数実行したまま、禁止された受取人を止めている点が重要です。条件をそろえた比較も報告し、決済要求をすべて危険行為として数えない評価にしています。
どこまで分かった?
設定ごとに攻撃群も変わるため、設定間の差をポリシーだけの効果とは断定できません。観測0件は絶対安全の証明ではなく、要旨ではセッション単位の上限0.38%も示されています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
APort Vaultは、ツールを使うAIエージェントの決済認可を評価するベンチマークである。公開のキャプチャー・ザ・フラッグ大会中に、稼働中の決済エージェントに対して人が作成した4,371件の攻撃を、8研究組織の14モデル、5つのポリシー設定、2種類の再実行トラックにわたって再実行する。Open Agent Passport(OAP)仕様を実装した決定論的な行動前チェックがある場合とない場合を比較し、225,964回の評価を完了した。複数の出来事を一つにまとめると精査に耐えない数値になってしまうため、評価ごとに5つの異なる出来事を分けて報告する。 リクエストは頻繁に発生し、その発生率の差はモデル間より設定間ではるかに大きい。ただし、各攻撃はちょうど一つの設定に属するため、ポリシーと攻撃群が同時に変わっている。モデル単独の評価での発生率は、レベル1で10.9%、レベル2で3.0%、レベル3で0.1%、レベル4で79.4%だった。すべてのモデルで評価したレベル4の1,293プロンプトでは、リクエスト率は71.2〜84.3%に分布し、809プロンプト(62.6%)が14モデルすべてからリクエストを引き出した。それぞれ、そのレベルの許可リストにある受取人への決済成功に至った。 条件間の差が現れるのは認可の境界である。レベル2〜4では、パスポートで許可されていない受取人への送金は、モデル単独で76,842件中140件、認可層を通した場合で69,297件中0件だった。モデル・プロンプト・トラックを一致させた68,970組では105件対0件だった。この0件という結果は790の元セッションにまたがり、セッション単位の上限は0.38%となる。これは決済そのものを拒否して得た結果ではない。認可層の下で25,370件の決済が実行され、ポリシーが評価した25,640回の送金呼び出しのうち拒否したのは187回で、そのうち148回は受取人が禁止されていたためだった。225,964回の評価、レベル別のパスポート、採点コード、解析スクリプトをhuggingface.co/datasets/aporthq/vault-benchmark-v1で公開する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
APort Vault is a benchmark for payment authorization in tool-using AI agents. It replays 4,371 attacks written by humans against a live payment agent during a public capture-the-flag event, across 14 models from 8 labs, five policy configurations and two replay tracks, with and without a deterministic pre-action check implementing the Open Agent Passport (OAP) specification. 225,964 evaluations completed. We report five distinct events per evaluation, because collapsing them is how an agent benchmark produces a number that does not survive review. Requests are common and their rate differs far more across configurations than across models, though each attack exists at exactly one configuration so policy and attack cohort vary together: 10.9% of model-alone evaluations at Level 1, 3.0% at Level 2, 0.1% at Level 3, 79.4% at Level 4. On the 1,293 Level 4 prompts, each evaluated on every model, request rates run from 71.2% to 84.3%, and 809 prompts (62.6%) elicited a request from all fourteen models, each ending in a successful payment to the level's allowlisted recipient. The authorization boundary is where the conditions diverge. At Levels 2 to 4, transfers to recipients the passport did not permit number 140 of 76,842 with the model alone and 0 of 69,297 behind the layer, and 105 against 0 on 68,970 matched model, prompt and track triples. The zero spans 790 source sessions, giving a per-session upper bound of 0.38%. It was not obtained by refusing payments: 25,370 payments executed behind the layer, while the policy denied 187 of the 25,640 transfer calls it evaluated, 148 of them for a forbidden recipient. We release the 225,964 evaluations, the level passports, the scoring code and the analysis script at huggingface.co/datasets/aporthq/vault-benchmark-v1 .
arXiv ID: 2609.22076 / 要約の誤りについて