複数ロボットのガウシアン地図を統合して言葉で物体を特定
CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding
この論文をやさしく読む
ひとことで言うと
複数ロボットの3D地図を統合し、質問したロボットから見た位置関係で物体を言葉に対応づけます。
何に役立つ?
自分が直接見ていない対象も他機の観測から探す、協調的な場面理解に使うことが考えられます。
この研究の面白いところ
幾何と意味の整合性を保って地図を合わせ、実世界の参照対象mIoUを52.6%から68.8%へ改善します。
どこまで分かった?
回転誤差2.58度から0.15度への改善はシミュレーション、mIoUの改善は実世界の評価です。部分的な地図重複がある設定を扱っています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
実体を持つロボットが指示表現に基づいてシーンを理解するには、指定された視点から、物体や関係を中心とする言語クエリを実際の対象に対応付ける必要がある。局所的な意味付きガウシアン地図は1台の観測内でこうした対応付けを支援できるが、協調環境では、独立に再構成した地図を位置合わせして融合した後も、この能力を保たなければならない。このとき、指示された対象や文脈上の目印が別のエージェントの観測に由来する場合でも、空間関係はクエリを発したロボットの視点から解釈する必要がある。 本研究では、この問題を、融合地図上での協調的な指示表現のガウシアン・グラウンディングとして定式化する。必要となるのは、幾何学的に位置合わせできること、インスタンス単位で意味を比較できること、そして視点を条件とした関係推論である。既存の言語対応ガウシアン手法は主に単一地図への問い合わせを扱い、ガウシアン位置合わせ手法は、言語による対応付けに必要な意味の互換性を保たず、幾何学的または測光的な位置合わせを最適化する。 そこで、協調的な指示表現対応ガウシアンスプラッティングの枠組みCoRef-GSを提案する。まず、語彙を固定せずインスタンスを識別する局所ガウシアン地図を構築する。次に、幾何学的・意味的な一貫性に基づくエージェント間位置合わせモジュールで、部分的に重なる地図を位置合わせし、視点を条件としたマスク関係グラフでクエリを対象に対応付ける。さらに、実世界とシミュレーションの屋内シーンを含む、2台の四脚ロボット用ベンチマークCoQuad-Refを導入する。実験では、シミュレーションシーンの回転誤差が粗い初期化後の2.58度から精密化後の0.15度へ低下し、実世界の指示対象に対するmIoUはReferSplatの52.6%から68.8%へ向上した。ベンチマークとソースコードは https://github.com/ruojiruoli17/CoRef-GS.git で公開する予定である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independently reconstructed maps are aligned and fused. In this setting, the referred target or its contextual landmark may come from another agent's observations, while spatial relations must still be interpreted from the querying robot's viewpoint. We formulate this problem as cooperative referring Gaussian grounding over fused maps, which requires geometric alignability, instance-level semantic comparability, and view-conditioned relation reasoning. Existing language-aware Gaussian methods mainly focus on single-map querying, whereas Gaussian registration methods optimize geometric or photometric alignment without preserving language-grounding-oriented semantic compatibility. We propose CoRef-GS, a cooperative referring Gaussian splatting framework. CoRef-GS constructs local open-vocabulary instance-aware Gaussian maps, then aligns partially overlapping maps with a cross-agent alignment module by geometric and semantic consistency, and grounds queries using a view-conditioned mask relation graph. We further introduce CoQuad-Ref, a dual-quadruped benchmark spanning both real-world and simulated indoor scenes. Experiments show that, on simulated scenes, CoRef-GS reduces the rotation error from 2.58{\deg} after coarse initialization to 0.15{\deg} after refinement, and improves real-world referring mIoU over ReferSplat from 52.6% to 68.8%. The established benchmark and source code will be publicly released at https://github.com/ruojiruoli17/CoRef-GS.git.
著者のコメント
The established benchmark and source code will be publicly released at https://github.com/ruojiruoli17/CoRef-GS.git
arXiv ID: 2609.20586 / 要約の誤りについて