指示に必要な目印だけを記録する視覚・言語ナビゲーション
SparseNav: Instruction-conditioned Sparse Semantic Perception for Training-Free Vision-Language Navigation
この論文をやさしく読む
ひとことで言うと
ロボットが文章の道案内を実行するとき、その時点で必要な目印だけを探して地図に記録する方法です。
何に役立つ?
無関係な物体の認識や記録を減らし、視覚・言語モデルによる経路選択を簡潔にする設計に役立ちます。四足歩行ロボットの屋内実機でも試しています。
この研究の面白いところ
目印が見え、しかもその位置が次の判断に役立つときだけ意味情報を取得する仕組みです。事前地図なしの実機運用も報告されています。
どこまで分かった?
成功率42.8%と40.7%はそれぞれ指定の未見環境検証分割での値です。実機での検証は複数の屋内環境に限られ、あらゆる環境での性能は要旨から分かりません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
地図を使う視覚・言語ナビゲーション(VLN)は、言語理解と幾何学的な経路計画をつなぐため、継続的な空間表現に頼る。しかし現在の指示に必要のない意味情報まで取得すると、知覚処理の費用と無関係な注釈が増える。関係のない物体を記録し続けることは計算を浪費するだけでなく、視覚・言語モデル(VLM)の計画器が使う視覚・空間表現を煩雑にする。この問題に対し、意味に基づくナビゲーションで必要な情報を絞る、追加学習不要の枠組みSparseNavを提案する。SparseNavは軽量な幾何学的鳥瞰図(BEV)地図と少数の目印の記憶を維持し、現在有効な部分指示を使って、位置付ける価値のある意味情報だけを必要に応じて取得する。まず指示管理器が進行状況を追い、現在探すべき目印を特定する。次に、求める目印が見え、その距離を含む位置が次の判断に役立つとき、指示条件付きの知覚機構が開放語彙の領域分割を実行する。得られた目印の記憶により、VLMは探索境界と局所方向を組み合わせた経由点候補から選べる。追加学習なしで、未見環境の検証分割においてR2R-CEで42.8%、RxR-CEで40.7%の成功率を達成した。制御した要素除去実験で、意味知覚の戦略と枠組みの各構成要素の寄与を調べた。さらに、幾何地図の作成と目印の位置付けにIntel RealSense D455 RGB-Dカメラを、自己位置推定にLivox MID-360 LiDARを備えたUnitree Go2四足歩行ロボットに、事前作成地図なしでSparseNavを導入した。複数の屋内環境で、指示に応じた経由点ナビゲーションの有効性を検証した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Map-based vision-language navigation (VLN) relies on persistent spatial representations to connect language understanding with geometric planning. However, acquiring semantics beyond the needs of the current instruction can introduce unnecessary perception cost and irrelevant annotations. Continuously accumulating unrelated objects may not only waste computation, but also clutter the visual-spatial representation consumed by the vision-language model (VLM) planner. To address this problem, we present SparseNav, a training-free framework that follows a less-is-more principle for semantic navigation. SparseNav persistently maintains a lightweight geometric bird's-eye-view (BEV) map and sparse landmark memory, acquiring new semantics on demand using the active sub-instruction to decide what is worth grounding. An instruction manager first tracks navigation progress and identifies the active landmark query. An instruction-conditioned perception mechanism then invokes open-vocabulary segmentation when the queried landmark is visible and its metric location can inform the next decision. The resulting landmark memory supports VLM selection among hybrid frontier and local directional waypoint candidates. Without any additional training, SparseNav achieves success rates of 42.8% on R2R-CE and 40.7% on RxR-CE, both on the Val-Unseen splits. Controlled ablations examine semantic perception strategies and the contributions of individual framework components. Furthermore, we successfully deployed SparseNav on a Unitree Go2 quadruped equipped with an Intel RealSense D455 RGB-D camera for geometric mapping and landmark grounding and a Livox MID-360 LiDAR for localization, without a prebuilt map. We validated its effectiveness across multiple indoor environments using instruction-conditioned waypoint navigation.
arXiv ID: 2609.26408 / 要約の誤りについて