arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

予測器を飛行時に使わない単眼カメラのドローン航法

Skytopia: Monocular Drone Navigation with Action-Conditioned Latent World Models

Yuhang Zhang, Rangya Zhang, Yujing Shang, Zhuoyuan Yu, Weiying Wang, Steven Yang, Qingsong Yan, Chao Yan, Mir Feroskhan

この論文をやさしく読む

ひとことで言うと

単眼カメラのドローンに世界モデルの表現を学ばせ、飛行時には重い予測器を使わずに航行する方法。

何に役立つ?

計算資源の限られるドローンで、位置や画像を目標とする航法を一つの方策で扱う設計に役立つ可能性がある。

この研究の面白いところ

学習時には次の観測を予測させるが、実行時には予測器を捨てても、シミュレーションでは三つの目標条件で比較手法を上回った。

どこまで分かった?

成功率57.8%、66.0%、49.0%はシミュレーションの各条件に対応する。実機でも目標到達を報告しているが、実機での成功率や長期運用の性能は要旨にない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

単眼カメラによるドローンの航法では、前方を向いた1台のカメラだけから未知の環境で目標に到達しなければならず、奥行きや大きさの手掛かりが少ない。世界モデルは行動に応じた観測の変化をモデル化するが、従来は飛行時にも予測を生成し、制御の各段階でその結果を行動生成に戻して使う。著者らは、方策が世界モデルから必要とするのは予測そのものではなく、予測を作るために必要な表現だと主張する。飛行中、実行した行動が観測間の変化の大部分を説明するため、予測は既知の移動量に基づく静的な場面の再投影に近くなる。そこで、行動を条件とする潜在世界モデルに基づく方策Skytopiaと、その学習に用いる3次元ガウシアン・スプラッティングの基盤を導入する。順方向の目的では意図した動きから次の観測の表現を予測し、逆方向の目的では予測された変化からその動きを復元する。予測自体は行動生成に渡さないため、予測器を取り除き、一つの方策を位置目標、画像目標、目標なしの三つの航法に利用する。シミュレーションでは、三つの条件すべてで比較手法を上回り、成功率はそれぞれ57.8%、66.0%、49.0%だった。予測器を取り除くと推論コストは59.4%減った。同じ方策を追加学習なしで実機ドローンに配備したところ、屋内、開けた屋外、森林の環境で目標に到達した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Monocular drone navigation requires reaching a goal in an unseen environment from a single forward-facing camera, which offers few cues for depth and scale. World models address this by modelling how observations evolve under actions, but they are built to be executed: the prediction is produced at deployment and fed back into action generation at every control step. We argue that what a policy needs from a world model is not the prediction but the representation required to produce it: in flight the executed action explains almost all of the change between observations, so prediction reduces to reprojecting a static scene under a known displacement. We therefore introduce skytopia, a policy built on an action-conditioned latent world model, and the 3D Gaussian Splatting platform on which it is trained. A forward objective predicts the representation of the next observation from the intended motion, and an inverse objective recovers that motion from the predicted transition. Because the prediction never reaches action generation, the predictor is discarded and one policy serves point-goal, image-goal, and goal-free navigation. Simulation experiments show that skytopia outperforms every baseline under all three specifications, attaining 57.8%, 66.0%, and 49.0% success rate, while discarding the predictor removes 59.4% of the inference cost. The same policy is subsequently deployed on a physical drone without fine-tuning and reaches goals in indoor, open outdoor, and woodland environments.

arXiv ID: 2609.26007 / 要約の誤りについて