arXiv論文メモ
新着一覧
cs.CR / cs.AI · 査読状況未確認

LLMエージェントのツール実行直前に権限と作用を検証

PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

Fengpeng Li, Qizhou Wang, Yuke Hu, Kemou Li, Jun Liu, Haiwei Wu, Jiantao Zhou, Di Wang

この論文をやさしく読む

ひとことで言うと

外部文書などに誘導されても、ツールを実行する直前に、その操作が利用者の認めた範囲に収まるかを検査する仕組みです。

何に役立つ?

考えられる用途は、検索やファイル操作などを行うエージェントの権限外動作を減らすことです。要旨では8つの安全性ベンチマークで攻撃成功率と有用性を評価しています。

この研究の面白いところ

入力を事前に審査するだけでは見分けられないケースを定式化し、実際に起こる作用の検証へつなげています。形式的な保証条件と、修復などを含む評価構成を区別している点も重要です。

どこまで分かった?

保証には最終動作が認証された経路切断を保持する条件があります。適応的探索は縮小規模の30件であり、成功0件はあらゆる攻撃への安全性を意味しません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ツールを使う大規模言語モデル(LLM)エージェントは、生成したテキストを現実の副作用へ変換する。そのため、汚染されたツールのメタデータ、取得ページ、メモリ、再利用可能なスキルが、次の呼び出しを誘導し得る。受け入れる前に成果物を審査するだけでは、この問題は解決しない。安全な版と情報を漏えいする版が同じ受け入れ審査の証拠を生むことがあり、その場合、健全なゲートはどちらについても当該箇所の制限を緩められない。本研究はこの条件を厳密に定式化し、配備側がなお介入できる最後の境界を明らかにする。 本研究では、すべてのツール呼び出しを実行直前に仲介するProvenance-Aware Capability Enforcement(PACE)を提案する。経路の閉じ込めは、表現された影響経路を実行時に切断する案を提示する。一方、ケイパビリティと作用の検証は、スキーマで定義された作用を、認証された要求から構成した権限と照合する。保証の対象となる実行契約と、評価に用いる構成とを区別する。評価構成では、ブロック案の後で権限のある呼び出しを復元したり、明示された修復を適用したりできる。閉じ込めの成立には、最終的な動作が保証対象の切断を保持する必要がある。 3系列の対象モデルを用いた8つの実行可能なエージェント安全性ベンチマークでは、評価構成の攻撃成功率は、比較対象となる79の攻撃列のうち62で単独最小、14で同率最小だった。各ベンチマーク本来の有用性指標を全体で評価した場合、防御なしのエージェントに対する低下は最大3ポイントだった。1,167の対応付き事例による完全なアブレーションでは、安全性向上の大部分が作用検証に、拒否の制御が境界での適応に由来すると示された。規模を縮小した適応的探索では、防御に対する権限外の標的30件中、成功は0件だった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking variant can produce the same admission evidence, and a sound gate then cannot relax that site for either. We make that condition precise, which leaves the last boundary a deployment can still act on. We present Provenance-Aware Capability Enforcement (PACE), which mediates every tool call immediately before it executes. Path confinement proposes an executable cut of represented influence paths, while capability and effect verification checks schema-defined effects against authority compiled from the authenticated request. We distinguish the certified execution contract from the evaluated configuration, which can restore an authorized call after a proposed block or apply a declared repair. Confinement requires the final action to preserve the certified cut. On eight executable agent-security benchmarks with three target-model families, the evaluated configuration gives strictly lowest attack success in 62 of 79 eligible attack columns and ties in 14; full-benchmark native utility loses at most three points relative to the undefended agent. A complete ablation over 1167 paired cases attributes most security gains to effect verification and refusal control to boundary adaptation. A reduced-scale adaptive search succeeds on 0/30 out-of-authority targets against the defense.

arXiv ID: 2610.01349 / 要約の誤りについて