arXiv論文メモ
新着一覧
cs.RO · 査読状況未確認

地形の物理的な負担を予測して野外ロボットを誘導

PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation

Aoran Jiao, Wenda Zhao, Hshmat Sahak, Timothy D. Barfoot

この論文をやさしく読む

ひとことで言うと

ロボットが形だけでは通れないと判断した場所について、画像と言語モデルで走行の負担も考え直す仕組みです。

何に役立つ?

想定用途は複雑な野外地形で人の介入を減らすことです。実際の混合地形ルートで自律走行率と介入間距離の改善を報告しています。

この研究の面白いところ

VLMの判断をそのまま使わず、エネルギー、振動、滑りの予測が実測とどれだけ合うかで重み付けします。必要な場面だけ再計画する設計も特徴です。

どこまで分かった?

評価は合計約6.4 km、五回の反復試行です。97.0%は報告された自律走行率であり、あらゆる地形での安全保証を意味しません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

地形評価はオフロード移動ロボットに不可欠な能力であり、非構造的で幾何学的に複雑な環境を安全かつ確実に走行することを可能にする。従来の幾何情報に基づく地形評価は計算が速いが、非構造的環境では保守的になりすぎることが多い。本研究では、野外ロボット向けに、従来の幾何情報による計画へ視覚言語モデル(VLM)の意味的推論を追加する、物理情報を取り入れた視覚言語オフロード走破性ナビゲーションシステムPIVOTを提示する。 評価を物理的な根拠に結び付けるため、VLMが予測する走行エネルギーコスト、ロボットの振動、車輪の滑りが、実測値とどの程度強く相関するかを定量化する。そして、各モダリティを予測と測定の相関で重み付けする統一的な走破性スコアを導入する。効率を保つため、通常は幾何情報による計画を用い、それで経路が見つからない場合にだけ意味的な再計画を呼び出す、二段階のナビゲーション構造を設計した。 複数の地形が混在する経路で、合計約6.4 kmにわたる閉ループ試行を五回繰り返した。幾何情報だけのナビゲーションと比べ、提案システムは全体の自律走行率を59.6%から97.0%へ高め、人の介入回数を11回から3回へ減らし、介入間の平均走行距離(MDBI)を69.2 mから412.9 mへ増加させた。これらの結果は、物理的な根拠を持つVLMによる地形評価が、効率的な幾何計画を通常モードとして保ちつつ、幾何情報だけの限界を越えて自律走行を大きく拡張できることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Terrain assessment is a critical capability for off-road mobile robots, enabling safe and reliable navigation through unstructured and geometrically complex environments. Conventional geometry-based terrain assessment is fast to compute but often overly conservative in unstructured environments. We present PIVOT: a Physically Informed Vision-Language Off-Road Traversability navigation system that augments conventional geometry-based planning with vision-language-model (VLM)-based semantic reasoning for field robots. To physically ground this assessment, we quantify how strongly the VLM's predicted traversal energy cost, robot vibration, and wheel slip correlate with real-world measurements and introduce a unified traversability score that weights each modality by its prediction-measurement correlation. For efficiency, we design a two-level navigation architecture that retains geometry-based planning as the nominal mode and invokes semantic replanning only when that mode fails to find a path. Across five repeated closed-loop trials on a mixed-terrain route totalling around $6.4$ km, the proposed system increases overall autonomy from $59.6\%$ to $97.0\%$, reduces human interventions from $11$ to $3$, and increases the mean distance between interventions (MDBI) from $69.2$ m to $412.9$ m compared with geometry-only navigation. These results demonstrate that physically grounded VLM-based terrain assessment can substantially extend autonomous navigation beyond the limitations of geometry alone, while preserving efficient geometric planning as the nominal mode.

arXiv ID: 2609.20983 / 要約の誤りについて