arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

複数AIの判断を整合させる共有状態と確定処理

Global Coherence: When Every Agent Is Right and the Team Is Still Wrong - A Local-to-Global Semantic Foundation for Multi-Agent Collaboration

Xin Heng

この論文をやさしく読む

ひとことで言うと

複数AIが協力するとき、個々の推論が正しくても、共有状態や変更履歴が欠けると全体として失敗する問題を扱います。

何に役立つ?

考えられる用途は、複数エージェントが共有予算や状態変更を扱うシステムで、確定時の検査と正本の管理を設計することです。

この研究の面白いところ

情報が見えない場合の成功率の限界と、共有状態を実行基盤が管理した場合の実験を結び付けています。モデルの強さと状態の完全性を切り分けます。

どこまで分かった?

不可能性定理は、同じ観測から許容行動を区別できない設定に対する結果です。実験は記載のベンチマーク条件に限られ、TeamBenchの数値は5試行です。要旨には九つすべての研究の詳細はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

AIエージェントは、それぞれ局所的に妥当な判断をしても、共同では不適切な結果を生み得る。これを大域的整合性問題と呼ぶ。これはモデルの知能だけではなく、共有状態の失敗である。本研究の観測別名化不可能性定理は、その厳密な境界を与える。ある方策が妥当な行動を保証できるのは、同じ観測を生むすべての世界に共通の許容行動がある場合、かつその場合に限る。区別できないk個の世界が互いに素な行動を必要とするなら、ランダム化した場合の最良の最悪時成功率は1/kであり、推論、役割、メッセージ、標本を増やしても欠けた区別は回復できない。より強いモデルは文脈内ではよく推論できても、その外を見ることはできない。 次に、局所から大域への実行時意味論X = (H,C,G,F;D)を与える。位相構造Hは重複する範囲を記録し、圏Cは状態を変える行動を管理し、亜群Gは可逆な変換を保持し、層Fは局所的な見方が一つの世界に貼り合わさるかを検査する。最小履歴Dは、将来の合法な行動を変える区別だけを残す。モデルは提案し、実行基盤が共有状態を保有して確定処理を管理する。 九つの研究で失敗とその境界を検証する。制御された改訂ベンチマークでは、同じ最先端モデルが、決定的な出来事を見られる場合は40/40を得る。見られない場合、評価した条件の得点は12〜17/40で、偶然の成功率1/3と整合する。正本に基づく事実を一つ復元すると40/40に戻る。TeamBenchでは、通常のチームは5試行中5試行で共有予算を超え、現在の件数を見せても5試行中4試行で違反し、確定時に制約を強制すると5試行中0試行となる。tau2-bench Telecomでは、通知されない巻き戻しの後、現在状態の検査は0.07、実行基盤は1.00を得る。通常のソルバーが関連する完全な状態を既に保有する場合には、予測どおり実行基盤と同点になる。直観に反する結論は、局所的な知能では、欠けた大域的状態を代替できないということである。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

AI agents can each make locally valid decisions yet jointly produce an invalid result. We call this the global coherence problem: a failure of shared state, not merely of model intelligence. Our Observation-Aliasing Impossibility Theorem gives the exact boundary. A policy can guarantee a valid action exactly when all worlds producing the same observation share an admissible action. If k indistinguishable worlds require pairwise-disjoint actions, the best randomized worst-case success is 1/k; more reasoning, roles, messages, or samples cannot recover the missing distinction. A stronger model can reason better within its context, but it cannot see beyond it. We then give local-to-global runtime semantics X = (H, C, G, F; D): topology H records overlapping scopes; category C governs state-changing actions; groupoid G retains reversible translations; sheaf F tests whether local views glue into one world; and minimal history D keeps only distinctions that alter legal futures. Models propose; the harness owns shared state and governs commit. Nine studies test both the failure and its boundary. On a controlled revision benchmark, the same frontier model scores 40/40 when the deciding event is visible; when it is hidden, tested arms score 12--17/40, consistent with chance (1/3); restoring one authoritative fact returns 40/40. On TeamBench, ordinary teams exceed a shared budget in 5/5 runs, a visible live count leaves 4/5 violations, and commit enforcement leaves 0/5. In tau2-bench Telecom, current-state checks score 0.07 after silent reverts, while the harness scores 1.00. Where a conventional solver already owns the complete relevant state, it ties the harness as predicted. The counterintuitive conclusion is that local intelligence cannot substitute for missing global state.

arXiv ID: 2610.02036 / 要約の誤りについて