arXiv論文メモ
新着一覧
cs.RO / cs.LG / cs.MA · 査読状況未確認

知覚特徴を共有して協調する複数ロボット

Latent Telepathy: Multi-Robot Communication with Self-Supervised Perceptual Latents

Howard Wang, Han Zheng, Cathy Wu

この論文をやさしく読む

ひとことで言うと

ロボット同士が、位置だけでなく見えている危険を圧縮された知覚ベクトルで伝える方法。

何に役立つ?

一台には見えない障害がある複数ロボットの協調行動に、通信量を抑えた知覚共有として役立つ可能性がある。

この研究の面白いところ

通信条件を固定した比較で、潜在ベクトルは見えない危険を99.7%回避し、186倍大きい生画像より信頼性が高かった。

どこまで分かった?

特定の課題設定でのシミュレーションと、物理ロボットのカメラからの102判断での復号結果である。現実の任意の危険や環境での回避率は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

部分的にしか環境を観測できない分散型の複数ロボットチームでは、あるロボットの次の行動を決める事実が、仲間にしか見えていないことが多い。既存の分散型手法が伝える位置や予定軌道などの運動情報では、仲間が何を知覚したかを伝えられない。複数エージェント強化学習による学習済み通信は知覚情報を運べるが、メッセージが課題に結びつき、中身も不透明になる。本研究はLatent Telepathyを提案する。 各ロボットは、自分の知覚のためにすでに計算している潜在ベクトルを送る。これは自己教師ありの共同埋め込み予測目標で学習したエンコーダーの出力であり、エンコーダーは固定され、チーム全体で共有される。受信側は課題の報酬だけから、そのベクトルに応じた行動を学ぶ。エンコーダーは知覚のためにもともと動くので、メッセージのための追加計算は不要で、通信量は一つの小さなベクトルで済む。方策を学習する前にエンコーダーを固定するため、メッセージは全ロボットに同じ意味を持つが、受信側にはその意味を直接教えない。 帯域、遅延、接続構造、受信者を固定し、メッセージの内容だけを変える比較手順で評価した。潜在ベクトルを送ると、ナビゲーターは見えない危険を99.7%のエピソードで回避し、雑音のない人手設計のメッセージと同等だった。位置や軌道のメッセージは偶然並みで、生のカメラ画像は通信量が186倍でも圧縮された潜在ベクトルより信頼性が低かった。結果は離散的な格子環境から、連続速度制御の画像環境まで成り立った。物理ロボットのカメラからも、102回の実際の判断全てで危険を復号した。また、複数エージェント強化学習の通信結果を連続制御へ移すには、メッセージが影響する判断へ探索によって到達できる必要があると特定し、その条件を回復する方法を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In a decentralized multi-robot team under partial observability, the fact that decides a robot's next action is often visible only to a teammate. Existing decentralized methods communicate kinematic information, such as position or planned trajectory, which cannot convey what the teammate perceives. Learned communication in multi-agent reinforcement learning (MARL) can carry perceptual content, but the resulting messages are task-coupled and opaque. We propose Latent Telepathy. Each robot broadcasts the perceptual latent vector it already computes for its own use, the output of an encoder trained with a self-supervised joint-embedding predictive objective, frozen, and shared across the team. A teammate learns to act on it from task reward alone. Because the encoder already runs for perception, the message costs no additional computation and a single compact vector of bandwidth. Because the encoder is frozen before any policy is trained, the message means the same thing to every robot, and the receiving robot is never told what it means. We evaluate Latent Telepathy with a content-controlled protocol in which bandwidth, latency, topology and receiver are held fixed and only the message content varies. Broadcasting the latent lets a navigator avoid an occluded hazard in 99.7% of episodes, matching a noiseless hand-designed message. Position and trajectory messages remain at chance, and the raw camera image, 186 times wider, is less reliable than the compressed latent. The result holds from a discrete gridworld to rendered pixels under continuous velocity control, and the encoder decodes the hazard from a physical robot's camera in 102 of 102 live decisions. We also identify a requirement for porting MARL communication results to continuous control, that the decision a message informs must remain reachable by exploration, and show how to restore it.

著者のコメント

8 pages, 6 figures. Submitted to IEEE ICRA 2027

arXiv ID: 2609.23269 / 要約の誤りについて