停滞した複数エージェントに対象を絞って探索を追加
Anchor and Perturb: Lazy Agent Remediation by Exploration Injection
この論文をやさしく読む
ひとことで言うと
協調が止まったエージェントだけに探索を促し、うまく動いている仲間の行動は保つ方法です。
何に役立つ?
複数エージェントの強化学習で協調が停滞した際の修復策として考えられる。
この研究の面白いところ
全員を同時に探索させず、低調なエージェントの特定の座標だけを変動させる。
どこまで分かった?
勝率の数値は要旨にあるテレメトリーのベンチマークでの結果であり、他の環境への一般化は示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Anchor and Perturb(AnP)は、探索時の変動の注入と、繰り返し学習で得た安定な解の維持を切り分け、複数エージェントの協調失敗を解消する軽量な枠組みである。既存の修復策は主に混合ネットワークの構造を変えるか、全エージェントに同時探索を強いるが、報酬が単調でない空間では時間差分学習に大きな不利益を生む。AnPは、働きが低下した「怠惰な」エージェントを特定し、その一部の座標へ非対称な探索の刺激を注入する。一方、収束した仲間は通常の貪欲な行動選択に固定する。実験で用いたテレメトリーのベンチマークでは、崩壊した共同方策を評価勝率5%の最低点から85%へ回復させ、最適でない協調状態からの脱出を助けた。ネットワーク構造を変更せずに、最高勝率90%を維持した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Anchor and Perturb (AnP) is a lightweight framework that resolves multi-agent coordination failures by decoupling exploratory variance injection from recurrent manifold stability. Existing remediation strategies predominantly alter mixing network architectures or enforce simultaneous exploration across the collective, which inevitably precipitates severe temporal-difference penalties in non-monotonic reward spaces. Specifically, AnP isolates underperforming lazy agents and injects an asymmetric exploratory pulse into targeted coordinates whilst anchoring converged teammates to nominal greedy exploitation. Empirical telemetry benchmarks demonstrate that AnP successfully rescues collapsed joint policies (recovering from a 5% evaluation win rate nadir back to 85%) and facilitates escape from suboptimal coordination plateaus, sustaining peak win rates of 90% without requiring structural network modifications.
著者のコメント
6 pages, 1 figure, work in progress
arXiv ID: 2609.27365 / 要約の誤りについて