arXiv論文メモ
新着一覧
cs.LG / cs.AI / cs.CL · 査読状況未確認

AI同士の推論中の情報共有は難問の発見を助けるか

Scaling Discovery through Test-Time Communication

Jongho Park, Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy, Dimitris Papailiopoulos

この論文をやさしく読む

ひとことで言うと

複数のAIが別々に問題を解く場合と、途中の発見を共有する場合を比較した研究です。難問では情報共有が有利になり得ますが、その効果は計算資源や評価方法に依存します。

何に役立つ?

考えられる用途は、明確な採点方法がある探索や最適化課題でのAIチーム設計です。通信を増やすだけでなく、十分な計算量と進捗フィードバックを用意する必要性を示します。

この研究の面白いところ

人数を増やしたときに情報共有の優位も広がる結果に加え、分類器圧縮では4体で1,957バイト・精度99.4%の提出物を得ています。単独試行の最良結果を選ぶ方式との比較が中心です。

どこまで分かった?

独立方式の方がよい条件も要旨に明記されています。ARC-AGI-3、充填、MNIST圧縮での成果であり、あらゆる研究作業や計算予算で通信が有利だとする結論ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

科学は孤立した作業ではなく協働によって進歩するが、既存のエージェントシステムはこの特徴をほとんど取り込めていない。エージェント間の通信が役立つかどうかは、先行研究の結果が分かれており、未解決の問いである。本研究は、突破口の共有が集団全体を前へ進められる難しい課題では、推論時の通信が、独立した並列試行を大きく上回り得ることを示す。 まず、新しい問題の解決を必要とするベンチマークARC-AGI-3で、複数エージェントによる推論時通信の規模を増やした効果を調べる。エージェントに事前の役割はなく、共有ディレクトリを通じて通信する。通信するk体のチームteam@kは、独立した4k体のエージェントと同じ成功率に達し、この優位はkとともに拡大した。これは規模拡大に伴って利得が積み重なることを示唆する。効果は効率にとどまらず、単独ではどのエージェントも解けない課題を、チームなら安定して解ける場合がある。 さらに、十分な計算資源があれば、この利得は研究的な課題にも移る。ポリオミノ充填では、通信するエージェントがbest@kを上回り、従来知られていた最良スコアを超えた。MNIST分類器の圧縮では、通信方式が既知の最良の人間による解を超えた。4体のチームが作成した分類器の提出物は1,957バイトで、テスト精度99.4%を達成し、既知の最良の人間による解と単独エージェントの最良結果の両方より小さかった。こうした利得は無条件ではない。計算資源が限られる場合や、進捗の明確な尺度がない場合には、独立したエージェントが通信方式を上回ることがある。しかし、十分な計算資源と明確なフィードバックの下では、複数エージェントの通信は一貫してより強い結果をもたらす。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Science advances not in isolation but through collaboration, yet existing agentic systems capture little of this. Whether communicating agents help remains an open question with mixed prior results. We show that test-time communication can substantially outperform independent parallel attempts on challenging tasks, where sharing a breakthrough can push the whole group forward. We first study the effect of scaling multi-agent test-time communication, where agents have no predefined roles and communicate via a shared directory, on ARC-AGI-3, a benchmark requiring novel problem solving. We find that a team of $k$ communicating agents, team@$k$, matches the success rate of $4k$ independent agents, and this advantage grows with $k$, suggesting gains compound with scale. The effect is not merely efficiency: a task that no single agent can solve, a team of agents can solve reliably. Furthermore, these gains transfer to research-oriented tasks, given sufficient compute. On polyomino packing, communicating agents outperform best@$k$ and exceed the prior best-known score. On MNIST classifier compression, communication surpasses the best-known human solution. A team of four agents produced a 1,957-byte classifier submission achieving 99.4% test accuracy, smaller than both the best-known human solution and the best single-agent result. These gains are not unconditional. Independent agents may outperform communication when compute is limited or when a clear measure of progress is absent. However, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.

著者のコメント

34 pages, 12 figures

arXiv ID: 2609.21032 / 要約の誤りについて