arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

ロボットの対象探索で場所と意味の不確かさを分ける

Spatial-Semantic Uncertainty in VLM-Based Target Search: Balancing Exploration and Identification

Alkesh K. Srivastava, Jonathan Diller, Vijay Kumar, Philip Dames

この論文をやさしく読む

ひとことで言うと

探す場所が分からないことと、見つけた物が目標か分からないことを分け、ロボットの探索を計画する研究です。

何に役立つ?

未発見の候補も考慮し、広く探すか候補を詳しく見るかを選ぶ仕組みです。観測が劣化した探索実験では、無作為探索より確信を伴う判断に到達しやすくなりました。

この研究の面白いところ

認識精度が似ていても、確信の妥当性は大きく違いました。意味の曖昧さを重視すると、不要な移動やVLMへの問い合わせを減らせる点が特徴です。

どこまで分かった?

500個の合成目標による評価と観測劣化条件での探索実験が根拠です。確信に達した割合は75.0〜92.5%で、無作為探索は20.0%でした。確信への到達率と正答率は別の指標です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

自然言語による記述を手掛かりに対象を探すロボットは、どこを探すかだけでなく、観測した候補のうちどれが目的の対象なのかも判断しなければならない。これらの判断には、候補の位置に関する空間的不確実性と、対象の同一性に関する意味的不確実性という、異なる二つの不確実性が関わる。しかし、視覚言語モデル(VLM)を用いる探索システムでは、両者がしばしば混同されている。本研究では、それぞれについて別々の信念分布を保持し、確率的なVLMの証拠を、未発見の対象に割り当てる確率も含めた全体的な対象同一性の事後分布へ統合する、空間・意味的不確実性の定式化を導入する。 この分解によって、情報理論に基づくプランナーは、空間的・意味的な期待情報利得(EIG)を通じて、候補の発見と対象の曖昧さの解消を個別に評価できる。これにより、広範な探索と早期の対象特定との間を調整する明示的な仕組みが得られる。500個の合成対象を用いて、VLMから不確実性を引き出す六つのインターフェースを評価した。その結果、認識精度が似ていても、確率の較正や誤った確信の程度には大きな違いが隠れていることが分かった。 観測を劣化させた探索・特定実験では、EIGに基づくプランナーが確信を伴う判断に至った試行の割合は75.0~92.5%であり、ランダム探索では20.0%だった。一方、いったん確信が得られると、空間と意味に対する重み付けが異なっていても、同程度の特定精度が得られた。意味に対する重みを増やすと、不要な探索とVLMへの問い合わせが減少した。これは、意味的不確実性を明示的に考慮して計画することで、判断の質を損なわずに対象の確定を速められることを示している。これらの結果は、身体を備えたVLMシステムにおける、不確実性の表現と、不確実性に基づく計画の役割がそれぞれ異なることを明らかにする。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Robots searching for a target from a natural-language description must determine not only where to search, but also which observed candidate is the desired target. These decisions reflect two distinct sources of uncertainty - spatial uncertainty over candidate locations and semantic uncertainty over target identity - that are often conflated in VLM-based search systems. We introduce a spatial-semantic uncertainty formulation that maintains separate beliefs over each component and integrates probabilistic VLM evidence into a global target-identity posterior, including probability mass for undiscovered targets. This decomposition allows an information-theoretic planner to independently value candidate discovery and target disambiguation through spatial and semantic expected information gain (EIG), providing an explicit mechanism for trading broader exploration against earlier identification. We evaluate six VLM uncertainty-elicitation interfaces on 500 synthetic targets and show that similar recognition accuracy can conceal substantial differences in calibration and false confidence. In degraded-observation search-and-identify experiments, EIG-based planners reach confident decisions in 75.0%-92.5% of trials, compared with 20.0% for Random search, while different spatial-semantic weightings achieve comparable identification accuracy once confidence is attained. Increasing semantic emphasis reduces unnecessary exploration and VLM queries, demonstrating that explicitly planning over semantic uncertainty can accelerate target resolution without sacrificing decision quality. These results highlight the distinct roles of uncertainty representation and uncertainty-driven planning in embodied VLM systems.

著者のコメント

Accepted for presentation at IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026) WorkshopRethinking Uncertainty for Modern Robotics Paradigms

arXiv ID: 2609.20443 / 要約の誤りについて