arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

検証結果に応じてAIエージェントの能力配分を調整

AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction

Md. Ashraful Babu

この論文をやさしく読む

ひとことで言うと

タスクに必要な機能を最初からすべて与える代わりに、検証で不足を診断して追加し、成功後は減らしても通るかを確かめる仕組みです。モデルや文脈、ツールなどを構成として調整します。

何に役立つ?

考えられる用途は、タスクを完了できる範囲で文脈量や利用可能なツール、推論資源を抑えることです。AB-ConfigBenchでは成功率と構成の縮小量を併せて評価しています。

この研究の面白いところ

失敗時の機能追加だけでなく、成功後に一つの次元を縮める反実仮想的な検証も行います。必要以上の能力が割り当てられていたかを、同じ検証結果が保たれるかで調べています。

どこまで分かった?

56.41%は1次元縮小の試行で検証結果を維持した割合で、全タスクの成功率ではありません。ツール数0は中央値です。異系列の追試では精度順位が再現せず、専用システムが有利な分野もあったと明記されています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルのエージェントは通常、あらかじめ決めた構成で配備される。しかし、必要なモデル能力、文脈、ツール、権限、記憶、計算資源は、タスクによって大きく変わり得る。本研究では、これらの要因を実行可能な構成として表し、検証に基づく診断、選択的な拡張、検証に基づく反実仮想的な縮小を通じて更新する、適応型AI Nano-AgentフレームワークAgentBettaを開発・評価する。評価では、制御された条件での機構検証と、外部エージェントとの比較を区別する。 AB-ConfigBenchでは、AgentBettaは検証済み成功率91.38%を達成した。同時に、すべての資源を与えた構成と比べ、文脈の割当量の中央値を64,000文字から8,000文字へ、利用可能にするツール数の中央値を5から0へ減らした。構成の不足を診断する機構は、評価した各次元にわたり適合率1.000、マクロF1スコア0.819を達成し、選択的な拡張によって無関係な構成次元の不要な変更を避けた。成功後の縮小では、評価した1次元だけの縮小試行の56.41%で検証結果が保たれた。これは、試した条件の下で、成功する構成の一部に除去可能な能力が含まれていたことを示す。 外部評価は、適応的な構成が、検証済みタスク完了と公開する能力の範囲とのバランスを改善し得ることを示す。ただし、結果はベンチマークとエージェントの系列によって異なる。とくに、異なる系列での追試では、主要な基盤モデルで得た正確さの順位が再現されず、特定分野では専用システムが優位を保った。以上は、AgentBettaを専用エージェント構造の普遍的な代替ではなく、能力配分と推論費用を調整する構成適応機構として解釈することを支持する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language model agents are typically deployed with predefined configurations, although the required model capability, context, tools, permissions, memory, and computational resources can vary substantially across tasks. This study develops and evaluates AgentBetta, an adaptive AI Nano-Agent framework that represents these factors as an executable configuration and updates them through verification-driven diagnosis, selective expansion, and verification-based counterfactual contraction. The evaluation distinguishes controlled mechanism validation from external agent comparisons. On the AB-ConfigBench benchmark, AgentBetta achieved 91.38% verified success while reducing median context allocation from 64,000 to 8,000 context characters and median tool exposure from five tools to zero compared with the fully provisioned configuration. The configuration-deficiency diagnosis achieved a macro-F1 score of 0.819 with precision of 1.000 across the evaluated dimensions, and selective expansion avoided unnecessary changes to unrelated configuration dimensions. Post-success contraction preserved verification outcomes in 56.41% of evaluated one-dimension contraction probes, indicating that some successful configurations contained removable capability under the tested conditions. External evaluations indicate that adaptive configuration can improve the balance between verified task completion and capability exposure; however, the results vary across benchmarks and agent families. In particular, the cross-family replication did not reproduce the primary-backbone accuracy ordering, and specialized systems remained advantageous for certain task domains. These results support interpreting AgentBetta as a configuration-adaptation mechanism that regulates capability allocation and inference expenditure rather than as a universal replacement for specialized agent architectures.

arXiv ID: 2609.23512 / 要約の誤りについて