arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

AIエージェントの性能と費用を検証可能な資格情報にする

LEGIT: Credentialing Protocol for Trustworthy AI Agent Marketplaces

Steve Drew and Jiayu Zhou

この論文をやさしく読む

ひとことで言うと

AIの成績表に、どの構成で、どれだけの予算を使って測ったのかを添え、署名付きで確認できるようにする仕組みです。

何に役立つ?

エージェントを選ぶ際に、成功率だけでなく費用や評価条件を比較する基盤として役立ちます。市場への適用は提案であり、市場全体の改善効果が示されたという意味ではありません。

この研究の面白いところ

単発の認証と過去の評判を同じ識別主体に結び付けます。同程度の成功でも費用が異なるという評価結果を、資格情報の設計に反映しています。

どこまで分かった?

評判の有用性はフィードバックの信頼性に依存します。評判操作の費用解析も特定のSybil攻撃モデルに基づき、あらゆる不正を防ぐ保証ではありません。要旨に具体的な費用差は記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

能力の異なるAIエージェントが、購入者のために専門的なタスクを自律的に完了するエージェント市場が登場しつつある。こうした市場の大きな課題は、購入者が、自分のタスクで最もよく機能するエージェントを容易には判断できないことである。報告されたベンチマークのスコアは、タスク、ソフトウェア、予算が異なると、検証や比較が難しい場合がある。 私たちは、認証、評判、および市場での割り当ての提案を結び付ける資格情報プロトコルLEGITを導入する。認証では、測定された品質と解決済みタスク1件当たりの費用を、署名付きの記録を通じて、エージェントの構成、タスク領域、評価予算、証拠にひも付ける。評判では、報告されたフィードバックの信頼性に依存するという条件の下で、過去のタスク結果の記録を同じ識別主体に結び付ける。購入者とエージェントは資格情報の記録を検証し、任意で用意される視覚的なプロフィールを確認できる。 評価により、観測されたタスク成功の程度が同様でも、エージェント構成によって費用に違いがあること、そして比較が評価予算に依存することが明らかになった。これらの結果は、性能測定を実際に試験した構成と資源制限に結び付けることを支持する。補完的な解析では、明示したSybil攻撃モデルの下で、評判操作に必要な預託金と手数料を定量化する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers. A major challenge of such marketplaces is that buyers cannot easily determine which agent will perform best on their tasks. Reported benchmark scores may be difficult to verify or compare across tasks, software, and budgets. We introduce LEGIT, a credentialing protocol connecting certification, reputation, and proposed marketplace allocation. Certification binds measured quality and cost per solved task to an agent configuration, task domain, evaluation budget, and evidence through a signed record. Reputation links records of past task outcomes to the same identity, subject to the reliability of the reported feedback. Buyers and agents can verify credential records and inspect optional visual profiles. Evaluations reveal cost differences between agent configurations with similar observed task success, and show that comparisons depend on the evaluation budget. These results support binding performance measurements to the tested configuration and resource limits. A complementary analysis quantifies the deposits and fees required for reputation manipulation under a stated Sybil attack model.

著者のコメント

28 pages, 7 figures, 11 tables

arXiv ID: 2609.21325 / 要約の誤りについて