通信への攻撃に耐える分散型マルチエージェント強化学習
Fully Byzantine-Resilient Multi-Agent Reinforcement Learning
この論文をやさしく読む
ひとことで言うと
複数の学習エージェント間の通信が攻撃されても、攻撃がない場合と同じ収束先を得る方法を提案した。
何に役立つ?
通信への攻撃に備える分散型強化学習の設計に役立つ可能性がある。対象は指定された通信攻撃と関数の条件に限られる。
この研究の面白いところ
2ホップ通信の冗長性を使い、時間変化するネットワークでも攻撃なしの場合と同じ極限点への収束を証明した。
どこまで分かった?
価値・報酬関数の線形パラメータ化、通信層に限る辺攻撃などを仮定する。複数ロボットの隊形制御での実演が示されるが、一般的な攻撃への耐性は要旨からは分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
各エージェントが局所的な相互作用を通じて共同で方策を学ぶ、分散型のビザンチン攻撃耐性を持つアクター・クリティック型マルチエージェント強化学習(AC-MARL)を研究する。従来の方法では、エージェントのパラメータが攻撃のない場合の極限点の近傍へ収束することしか保証されず、性能が低下する。本稿は、各エージェントが2ホップ先までのメッセージの冗長性を利用して信頼できるメッセージを見分ける分散型手法、Fully Resilient AC-MARL(FRAC-MARL)を提案する。価値関数とチーム報酬関数を線形にパラメータ化し、攻撃者の行動が通信層に限られるビザンチン型の辺攻撃を想定する。この条件の下、時間とともに変わる通信グラフでも、エージェントのパラメータが攻撃のない場合と同じ極限点にほぼ確実に収束することを証明する。また、収束のための新しいグラフ構造上の条件、その条件を満たすネットワークの体系的な構築方法、および条件を多項式時間で検証できることを示す。最後に、協調する複数ロボットの隊形制御課題で手法を実演する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We study distributed Byzantine-resilient actor-critic multi-agent reinforcement learning (AC-MARL), where agents collectively learn policies through local interactions. Existing methods guarantee convergence of the agents' parameters only to a neighborhood of the attack-free limit points, resulting in degraded performance. We propose Fully Resilient AC-MARL (FRAC-MARL), a decentralized method in which each agent leverages redundancy in two-hop messages to identify reliable messages. Under linear parameterizations of the value and team-reward functions and Byzantine edge attacks, where adversarial behavior is confined to the communication layer, we prove that agents' parameters converge almost surely to the same limit points as in the attack-free case over time-varying communication graphs. We introduce a novel topological condition for the convergence of our method, present a systematic method to construct such networks, and prove that this condition can be verified in polynomial time. Finally, we demonstrate our method on cooperative multi-robot formation control tasks.
著者のコメント
12 pages, 3 figures
arXiv ID: 2609.25701 / 要約の誤りについて