地図と観測の構造化記録でロボットの言語指示探索を支える
Structured World-State Reasoning for Agentic Robotic Search
この論文をやさしく読む
ひとことで言うと
言葉で指定された物を探す際、地図と新しい観測を更新されるグラフにまとめ、複数の解釈を追加観測で絞り込むロボット探索の研究です。
何に役立つ?
地図だけでは対象が分からない場面で、次にどこを見るべきかを計画するために役立ちます。CityNavでの比較に加え、実機ドローンで地図にない車両も含めた対象の特定を示しています。
この研究の面白いところ
推論、観測の検査、最終判断の役割を分け、分からなければもう一度観測します。同じモデルや移動予算にそろえた1,000例の比較もあり、構成の効果を調べています。
どこまで分かった?
全5,311例と共通1,000例の結果は評価集合が異なります。15.7や18.8は相対的な改善率ではなくパーセントポイント差です。実機で示したのは3つの言語対象で、広範な実環境での成功率ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
長い時間範囲にわたるロボットの探索では、文章情報、事前地図、時間とともに届く観測という、異種で不完全、かつ曖昧なことも多い根拠に照らして自然言語を解釈しなければならない。中心的な課題は、これらの情報の流れを文脈に位置付け、対象を選ぶ前にどこで根拠を集めるかを決めることである。本研究は、地理空間の事前情報で初期化し、知覚によって更新する持続的なグラフに推論を結び付ける枠組み、WORLDS(World-state Observation and Reasoning for Language-guided Discovery and Search)を提示する。並列のReasonerが競合する解釈候補を維持し、それらを区別するために必要な根拠を要求する。要求された観測を収集してマルチモーダルのExaminerで処理した後、Judgeが根拠に結び付いた対象を選ぶか、もう一度の処理を求める。 WORLDSはCityNavの全5,311テストエピソードで、報告済みで最高となる51.8%のナビゲーション成功率を達成した。OSMのみを事前情報に使う高解像度の正射投影プロトコルの下で、既公表の最高値を15.7パーセントポイント上回る。共通の1,000エピソードでは、同じモデル、事前情報、センシング構成、移動予算を使う最も強い適応済みベースラインの27.9%に対して、50.0%を達成する。Examinerによる観測ベースの検証は、この成功率に5.9ポイント寄与する。推論量を減らした設定でも、WORLDSは適応済みGeoNavベースラインを18.8ポイント上回り、生成トークン数も少ない。さらに、生成した観測用ウェイポイントを飛行するクアッドローターで実演し、地図に存在しない車両を含む3つの言語指定対象を、搭載カメラの画像から特定する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Long-horizon robotic search must resolve natural language against heterogeneous, incomplete, and often ambiguous evidence: textual information, prior maps, and observations arriving over time. The core challenge is to contextualize these streams and decide where to gather evidence before selecting a target. We present WORLDS: World-state Observation and Reasoning for Language-guided Discovery and Search, a framework that grounds reasoning in a persistent graph initialized from geospatial priors and updated by perception. Parallel Reasoners maintain competing candidate interpretations and request evidence to distinguish between them. We collect and process the requested observations with a multimodal Examiner, after which a Judge selects a grounded target or requests another pass. WORLDS achieves 51.8% navigation success across all 5,311 CityNav test episodes, the highest reported success rate, exceeding the previous published best by 15.7 percentage points under an OSM-only, high-resolution orthographic protocol. On 1,000 shared episodes, it achieves 50.0% versus 27.9% for the strongest adapted baseline using the same model, prior, sensing stack, and movement budget. Observation-based verification by the Examiner contributes 5.9 points of this success, and at a reduced reasoning-effort setting WORLDS still exceeds the adapted GeoNav baseline by 18.8 points while generating fewer tokens. We also demonstrate WORLDS on a quadrotor, which flies the generated sensing waypoints and grounds three language targets, including a vehicle absent from the map, from its onboard imagery.
著者のコメント
9 pages, 5 figures, Submitted to the IEEE International Conference on Robotics and Automation (ICRA) 2027
arXiv ID: 2609.23841 / 要約の誤りについて