意味を持つメッセージで複数ロボットの作業を調整
DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication
この論文をやさしく読む
ひとことで言うと
各ロボットが自分の観測と相手からのメッセージを使って計画し、別の行動モデルで細かな動作を実行する分散型の協調方法です。
何に役立つ?
複数台で長い操作作業を進めるシステムの設計・評価に役立ちます。分散制御で協調が必要なRoboPolyという課題集も作っています。
この研究の面白いところ
全体を一つの制御器に集めず、各機体に高レベルの調整役と低レベルの実行役を置きます。異なる事前学習モデルの得意分野を使い分ける構成です。
どこまで分かった?
改善はRoboPolyとRoboTwinで評価されています。要旨には成功率の具体値、通信条件、台数を増やした場合の性能、実機評価の内訳は記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚言語モデル(VLM)と視覚言語行動モデル(VLA)は近年、汎用ロボットを急速に進歩させてきたが、その多くは単体ロボットの設定に集中している。複数ロボットへ拡張することは、信頼できる細かな動作を維持しながら長時間の行動を協調させる必要があるため、依然として難しい。私たちは、意味的なコミュニケーションによって複数ロボットを協調させる、分散型の階層フレームワークDuoMindを提案する。各ロボットは、低レベル実行にVLAベースの行動モデルを、高レベル推論とエージェント間の調整にVLMベースの調整役を用いる。 各計画ステップで、それぞれのロボットの調整役は、課題指示、局所観測、他ロボットから受け取ったメッセージをもとに推論する。その後、行動モデルへの低レベル指示と、相手ロボットへの意味的メッセージを生成する。この構成は、VLMの意味推論能力とVLAの精密な行動生成能力を組み合わせ、事前学習済みモデルの相補的な強みを利用する。 複数ロボット協調のベンチマーク不足に対処するため、分散制御のもとで協調的な閉ループ実行を必要とする長時間の操作課題からなるRoboPolyも開発する。RoboPolyとRoboTwinでの実験は、DuoMindが複数ロボットの課題性能を改善することを示し、要素除去実験は階層的な調整と意味的コミュニケーションの寄与を確認する。詳細はプロジェクトページで公開している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors while maintaining reliable, fine-grained execution. We introduce DuoMind, a distributed hierarchical framework for multi-robot coordination through semantic communication. Each robot uses a VLA-based action model for low-level execution and a VLM-based orchestrator for high-level reasoning and inter-agent coordination. At each planning step, the orchestrator at each robot reasons over the task instruction, local observations, and messages received from other robots. It then generates low-level instructions for the action model and semantic messages for peer robots. This architecture exploits the complementary strengths of pretrained models by combining the semantic reasoning capabilities of VLMs with the precise action-generation capabilities of VLAs. To address the scarcity of benchmarks for multi-robot coordination, we further develop RoboPoly, a benchmark comprising long-horizon manipulation tasks that require coordinated, closed-loop execution under distributed control. Experiments on RoboPoly and RoboTwin demonstrate that DuoMind improves multi-robot task performance, while ablation studies confirm the contributions of hierarchical orchestration and semantic communication. More details are available on our project page.
arXiv ID: 2610.02161 / 要約の誤りについて