生成動画と実測形状で地上・空中両用ロボットを誘導する
TADreamer: Zero-Shot Language-Guided 3D Navigation for Terrestrial-Aerial Bimodal Robots via Video Imagination
この論文をやさしく読む
ひとことで言うと
歩行・走行と飛行を切り替えるロボットについて、生成動画で想像した経路を実測の3D形状へ合わせて実行する方法です。
何に役立つ?
言葉の指示に応じて移動経路と移動様式を選ぶ際、生成映像の尺度のずれを補正する用途があります。専用の追加学習なしで構成しています。
この研究の面白いところ
視野条件による初期尺度推定と、実測点群への登録を二段階で行います。七つの屋内外場面で、各回五候補なら二回以内に使える生成動画を得ました。
どこまで分かった?
実環境七場面での検証です。深度誤差の87.7%・86.3%削減は較正に使った観測でのNavDreamer比較であり、未知環境での走行成功率とは区別する必要があります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
地上走行と飛行を切り替える両用ロボットの言語誘導ナビゲーションでは、場面の状況と課題の意図に合う経路と移動モードを選ぶ必要がある。生成動画はそのような動作系列を表現できるが、スケールの曖昧さと軸ごとに異なる幾何学的ひずみにより、距離尺度の整合したナビゲーション参照を復元することは難しい。 本研究では、課題固有の学習や微調整なしに、動画で想像したナビゲーションを実測形状に結び付けるゼロショットのフレームワークTADreamerを提示する。視覚言語モデルは、搭載センサーの観測と指示をナビゲーション用プロンプトへ変換し、有効な生成動画を選び、再生成が必要な場合には修正フィードバックを与える。選択した動画から、地上または空中のモードラベルを付けた3次元経由点を再構成する。2段階の較正では、まず視野の制約からスケール推定を初期化し、次に再構成点群を実測形状に位置合わせして、軸別スケール、回転、並進を精密化する。較正後の経由点とモードラベルを、実測形状を取り込むプランナーに与え、ロボットを動作させる。 実環境での実験では、屋内外7つのシナリオでナビゲーションを実証した。各ラウンド5候補を生成した場合、7シナリオすべてで2ラウンド以内に使用可能な動画が得られた。較正に用いた観測では、NavDreamerと比べ、平均絶対深度誤差を87.7%、平均絶対相対深度誤差を86.3%削減した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Language-guided navigation for terrestrial-aerial bimodal robots requires selecting routes and locomotion modes that match scene context and task intent. Generated videos can represent such motion sequences, but recovering metrically consistent navigation references from them is challenging because of scale ambiguity and axis-dependent geometric distortions. We present TADreamer, a zero-shot framework that grounds video-imagined navigation in measured geometry without task-specific training or fine-tuning. A vision-language model translates onboard observations and instructions into navigation prompts, selects valid generated videos, and provides corrective feedback when regeneration is needed. The selected video is reconstructed into 3D waypoints annotated with terrestrial or aerial modes. A two-stage calibration procedure uses field-of-view constraints to initialize scale estimation, then refines axis-dependent scales, rotation, and translation by registering the reconstructed point cloud to measured geometry. The calibrated waypoints and mode labels guide a planner that incorporates measured geometry for robot execution. Real-world experiments demonstrate navigation across seven indoor and outdoor scenarios. With five candidates per round, usable videos are obtained within two rounds in all seven scenarios. On the calibration observations, our method reduces mean absolute depth error by 87.7% and mean absolute relative depth error by 86.3% compared with NavDreamer.
arXiv ID: 2609.19824 / 要約の誤りについて