arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

空間変換器で多数のAIロボットの編隊を制御

Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers

Frederic Vatnsdal, Roshan Gopal, Romina Garcia Camargo, Vijay Kumar, and Alejandro Ribeiro

この論文をやさしく読む

ひとことで言うと

各ロボットが短い学習済みフィードバックを交換して、大きな群れの飛行を自然言語で制御する方法を研究した。

何に役立つ?

多数のロボットを分散的に制御し、言語指令と編隊維持を両立する方式の検討に役立つ可能性がある。

この研究の面白いところ

学習時の最大16倍にあたる最大1,024台まで拡張し、曖昧な未見の指令にも追加学習なしで対応した。生の状態を言語に入れる方法ではまとまりが崩れた。

どこまで分かった?

要旨は実験での群飛行結果を述べるが、実機かシミュレーションか、環境条件の詳細は記していない。現実の全条件への適用を示すものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)はロボットの計画や航行に新しい方法をもたらすが、チームが大きくなると単純な複数ロボット課題でも失敗する。本研究は、推論空間でのフィードバック制御を使い、多数のエージェント型ロボットの集団を制御する、拡張可能で分散型の構造COMPASSを提案する。各ロボットで空間変換器が複数ホップにわたる機体間のメッセージを集約し、学習したフィードバックトークンを作る。 実験では、言語モデルの集団は入力指令に構造的な多様性を持たせることで、偏りを打ち消し、規模を拡大しても性能向上を保てた。中央集権型の最先端LLM方策や、言語だけによる通信へ置き換えた比較条件と比べると、COMPASSの組み合わせ設計は、まとまりのある群飛行の編隊を作り、指令の意図どおり正確に飛行した。推論へのフィードバックは、短い学習済みトークンと組み合わせた場合に最もよく働いた。対照実験では、生の状態を言語チャネルに入れる手作業設計のフィードバックは、編隊のまとまりを失わせた。COMPASSは意味が曖昧な未見の指示にも追加学習なしで対応し、学習時の最大16倍、最大1,024台のロボットの群れを自然言語の指令で飛行させた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large Language Models (LLMs) introduce an exciting new paradigm for planning and navigation in robotics, but fail on even simple multi-robot tasks as team sizes grow. We propose COMPASS, a scalable, decentralized multi-robot architecture for controlling large collectives of agentic robots with reasoning space feedback control. Feedback is generated locally on each robot by a spatial transformer which aggregates multi-hop messages across the fleet into a learned feedback token. Our experiments find that collectives of language models demonstrate performance gains from structured diversity of the input command, which can cancel biases; an advantage that is held across scale. Compared against a centralized frontier LLM policy and a language-only communication ablation, we find that the coupled design of COMPASS decisively produces cohesive flocking formations that accurately fly the commanded intent. We show that reasoning feedback works best when composed with a compact learned token. Our ablations show that hand engineered feedback with raw state appearing in the language channel obliterates cohesion. COMPASS generalizes zero-shot to unseen instructions of ambiguous meaning while commanding flocks up to 16 times its training scale, flying up to 1024 robots under natural language commands.

arXiv ID: 2609.28247 / 要約の誤りについて