川を航行する無人艇のための言語指示と視覚の連携
RiverVLN: Phase-Grounded Temporal Vision--Language Navigation for Unmanned Surface Vehicles
この論文をやさしく読む
ひとことで言うと
川を進む無人艇が長い言語指示を小さな段階に分け、目印を確認しながら航行する方法を作った。
何に役立つ?
言語指示に従う無人水上艇の航行や、進行状況を確認しながらの再計画に役立つ可能性がある。
この研究の面白いところ
連続運動と慣性のある艇に対し、意味上の段階、視覚情報、移動履歴を一緒に使う。
どこまで分かった?
Unity–ROSでの成功率0.79と実機試験が報告されるが、すべての河川や気象条件での性能は示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚と言語を用いるナビゲーションは主に屋内や陸上ロボット向けに発展し、言語を静的な目標、移動を離散的でほぼ瞬間的な動作とみなすことが多かった。しかし、川の無人水上艇では慣性と操舵の制約のもとで連続的に動き、長い指示を、数が少なく見た目も曖昧な水上の目印を頼りに実行する必要がある。RiverVLNは、こうした長距離の川での連続運動を対象としたベンチマークである。PGT-NAVは、指示全体を直接動作へ変える代わりに、視覚的に確認できる意味上の段階を順番に並べ、画像と運動の証拠から現在の段階を更新する。現在の進行段階を視覚と運動の履歴、および段階ごとの目印の対応に融合し、局所的なSE(2)姿勢の増分を六つ予測する。艇は予測した軌道のW3へ向けて進み、再観測して段階と目印を更新し、地図に基づく安全層を通じて再計画する。この予測・実行・再観測のループによって、GNM型およびViNT型の比較法に対し位置と向きの誤差の蓄積を大幅に減らし、Unity–ROSでの閉ループ航行の平均成功率は0.79だった。未見の橋の開口部での試行と実物の無人艇での実験も行い、段階に基づく表現が制御された評価から実機へ移ることを示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-language navigation (VLN) has largely been developed for indoor and terrestrial robots, where language can often be treated as a static goal and motion is approximated by discrete or near-instantaneous actions. These assumptions break down for unmanned surface vehicles (USVs): river navigation requires continuous motion under inertia and limited maneuverability, while long-horizon instructions must be executed through sparse and visually ambiguous maritime landmarks. We introduce RiverVLN, to our knowledge the first benchmark designed for long-horizon USV VLN under continuous riverine motion, and PGT-NAV, a phase-grounded temporal navigation framework for USVs. Rather than directly mapping an entire instruction to motion, PGT-NAV converts it into an ordered sequence of visually verifiable semantic phases and maintains the active phase online through grounded visual and motion evidence. This explicit semantic progress state is fused with visual-motion history and phase-specific grounding to predict six local SE(2) pose increments. The resulting trajectory is executed in a predict-execute-re-observe loop, where the vessel executes toward W3, updates phase and grounding, and replans through a map-based safety layer. Experiments show that PGT-NAV substantially reduces recursive position and heading drift relative to GNM-style and ViNT-style baselines and achieves an average success rate of 0.79 in Unity-ROS closed-loop navigation. Unseen bridge-opening trials and real-world USV experiments further demonstrate that the phase-grounded representation transfers from controlled evaluation to physical USV deployment.
arXiv ID: 2609.23423 / 要約の誤りについて