複数AIの連携構造を一要素ずつ変えて診断する評価基盤
OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems
この論文をやさしく読む
ひとことで言うと
複数のAIを連携させた仕組みを、単なる最終点数ではなく、連絡方法や役割、誤情報への耐性を一つずつ変えて調べる評価基盤です。
何に役立つ?
あるシステムの性能がどの構成要素に支えられているかや、故障・誤情報に弱い箇所を診断するのに役立ちます。正解率とトークン効率を分けた比較もできます。
この研究の面白いところ
元の成績が同程度でも、誤ったメッセージや担当の停止への反応が異なることを示しています。タスクなどを固定して一要素だけを変えるため、最終スコアだけでは見えない違いを調べられます。
どこまで分かった?
結果は17構成、6分野29データセットと追加の400タスクに基づきます。専門担当の除去による影響が平均で大きいという結果は、すべてのタスクで批評担当より重要だとする結論ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
グラフを用いたマルチエージェントシステム(G-MAS)は、通信グラフと役割の割り当てを通じて大規模言語モデルのエージェントを協調させ、それらが情報交換や責任分担の方法を定める。しかし、システム間の最終スコア比較では、モデル、通信パターン、役割、計算コストの違いが混ざり、性能差を特定の通信構造、役割の割り当て、情報の流れに帰属させにくい。 この評価上の帰属問題に対処するため、制御された介入を通じて、これらの構成要素がG-MASの性能に与える影響を診断するベンチマークOpenMAS-GComを導入する。システムを、協働単位、通信リンク、共有される中間情報、実行規則によって表現する。OpenMAS-GComは、タスク、モデル、プロンプト、予算上限を固定したまま、構成要素を1つだけ変更した版と元のシステムを比較する。具体的には、通信辺のつなぎ替え、専門担当または批評担当エージェントの除去、中間メッセージの誤った内容への置換、実行中のワーカーの停止を行う。 ベンチマークでは、単一エージェント、通常のマルチエージェント、グラフを用いる構成の計17種類を、6分野の29データセットで評価する。さらに、複数文書の情報を組み合わせ、矛盾する記録を解決し、指定された値を出典識別子とともに返すことを求めるG-MAS-Complexタスク400件を追加する。実験では、批評担当を除去した場合より専門担当を除去した場合の方が平均的な性能低下が大きいこと、元のスコアが似ていても誤ったメッセージとワーカー停止に対する性能低下は異なること、G-MAS-Complexでは正解率が最も高い構成とトークン当たりの正解率が最も高い構成が異なることが示された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Graph-enhanced multi-agent systems (G-MAS) coordinate large language model agents through communication graphs and role assignments, which determine how agents exchange information and divide responsibilities. However, final-score comparisons across systems combine differences in models, communication patterns, roles, and computation costs, making performance differences difficult to attribute to specific communication structures, role assignments, and information flows. To address this evaluation attribution problem, we introduce OpenMAS-GCom, a benchmark for diagnosing how these components affect G-MAS performance through controlled interventions. We represent systems through collaboration units, communication links, shared intermediate information, and execution rules. OpenMAS-GCom compares original systems with versions modified by changing one component while keeping tasks, models, prompts, and budget limits fixed. We rewire communication edges, remove specialist or critic agents, replace intermediate messages with incorrect content, and disable workers during execution. The benchmark evaluates 17 single-agent, ordinary multi-agent, and graph-enhanced configurations on 29 datasets across six domains. We add 400 G-MAS-Complex tasks requiring agents to combine information from multiple documents, resolve conflicting records, and return specified values with source identifiers. Experiments show larger mean losses after specialist removal than after critic removal, different performance degradation under incorrect messages and worker failures despite similar original scores, and different configurations achieving the highest accuracy and accuracy per token on G-MAS-Complex.
arXiv ID: 2609.21527 / 要約の誤りについて