arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

将来予測を経路計画に直接使う自動運転モデルForeDrive

ForeDrive: Foresight-Guided End-to-End Autonomous Driving with a Planning-Relevant Latent World Model

Sinuo Wang, Zichong Gu, Yuhan Huang, Wenxin Wen, Xun Yang, Yiqing Zhang, Xingyu Zhang, Ningyu Che, Jie Ling, Qiankun Yu, Wei Liu, Jing Xu, Xinggang Wang

この論文をやさしく読む

ひとことで言うと

自動運転の経路生成で、将来予測の表現を計画器の入力として使うモデルです。

何に役立つ?

画像から経路を作るモデルに未来予測を組み込む設計の参考になります。要旨ではNAVSIM v1とv2での評価結果を示しています。

この研究の面白いところ

未来の予測をそのまま信用せず、時間幅ごとの信頼性と画像・軌道の位置ずれを考慮して、現在の観測と融合しています。

どこまで分かった?

推論時は現在の前方画像だけを視覚入力に使っています。要旨に実車走行試験の結果はありません。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

従来の潜在世界モデルは通常、未来を予測しやすいことを目標に最適化されるが、得られる表現が自動運転の計画に役立つとは限らない。予測は経路生成を直接条件づける信号ではなく、事前学習や補助的な教師信号に使われることが多い。本研究は、計画に関連する潜在表現を学習し、それをDiffusion Transformerの計画器に非対称に結びつけるForeDriveを提案する。計画器はJEPA型世界モデルで学習した複数の時間幅にわたる未来の潜在表現を受け取る。計画の勾配は共有のオンラインエンコーダーを更新する一方、勾配を止める経路により潜在予測器は予測損失だけで学習する。予測した未来の信頼性は時間幅によって異なり、鳥瞰視点の軌道と画像トークンは位置がそろわないため、ゲート付きの視覚融合、未来状態の注入、軌道適応バイアスTABを用い、現在の観測を上書きせずに未来の潜在表現を誘導情報として取り込む。純粋な模倣学習で学習し、推論時の視覚入力には現在の前方画像だけを用いる条件で、ForeDriveはNAVSIM v1でPDMS 89.9、NAVSIM v2で一段階のEPDMS 90.0を達成した。強化学習も外部の軌道評価器も使っていない。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-23 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Existing latent world models are typically optimized for future predictability, yet the resulting representations are not necessarily useful for planning in autonomous driving. Predictions are commonly used for pretraining or auxiliary supervision rather than as direct conditioning signals for trajectory generation. We propose ForeDrive, which learns a planning-relevant latent representation and couples it asymmetrically to a Diffusion Transformer (DiT) planner. The planner consumes multi-horizon latent future representations learned with a JEPA-style world model; planning gradients update the shared online encoder, while stop-gradient routing trains the latent predictor with forecasting losses only. Because predicted futures have varying reliability across horizons and BEV trajectories are misaligned with image tokens, we use gated visual fusion, future-status injection, and Trajectory-Adaptive Bias (TAB) to inject future latents as guidance without overriding the current observation. Trained with pure imitation learning and using only the current front-view image as visual input at inference, ForeDrive attains 89.9 PDMS on NAVSIM v1 and 90.0 one-stage EPDMS on NAVSIM v2, without reinforcement learning or an external trajectory scorer.

著者のコメント

9 pages, 4 figures; 8 pages supplementary with 4 figures

arXiv ID: 2609.26299 / 要約の誤りについて