少数の1次元トークンで車載画像から高精度地図を作る
MapLightning: Online Vectorized HD Map Construction with 1D Map Tokens
この論文をやさしく読む
ひとことで言うと
車載カメラの画像から道路の高精度地図を作る際、中間データを小さくして速度と精度を改善する方法です。
何に役立つ?
自動運転向けの地図生成や、その地図を使う軌道予測が用途です。2つのデータセットで精度・推論速度・メモリ使用量を比較しています。
この研究の面白いところ
鳥瞰図の密な格子を使わず、画像と少数の地図トークンをまとめて自己注意で処理します。表現を小さくしたことで、復号時には広い範囲の関係を扱えるようにしています。
どこまで分かった?
要旨の40 FPS超などは報告された比較条件での結果です。具体的な計算機構成や実走行時の安全性の検証は要旨に記載されていません。コード公開は予定として述べられています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
オンラインでのベクトル形式の高精度地図構築は、安全な自動運転を大規模に展開するために不可欠であり、正確かつリアルタイムの推論を必要とする。従来手法は一般に、中間表現として密な鳥瞰図(BEV)格子に依存している。私たちは、この密なBEV格子を、少数の学習可能な1次元地図トークンに置き換えるMapLightningを提案する。画像特徴から地図トークンを構成する際、通常の交差注意ではなく自己注意を選択する。これは、画像トークンと地図トークンの相互作用と文脈の集約を一緒に行えるためである。Transformerに基づく地図生成器は、地図と画像のトークンを連結して全面的な自己注意を適用し、画像トークンを捨て、更新された地図トークンを復号用に保持する。 この設計には3つの利点がある。第1に、トークン数とメモリ使用量が少なく、高速に動作する効率的な表現である。第2に、軽量な設計によって、地図デコーダーが変形可能な交差注意ではなく全面的な交差注意を使え、全体の文脈をよりよく扱える。第3に、BEVに基づく手法と異なり、ネットワークはカメラの投影パラメータを使わないため、カメラ外部パラメータの摂動に頑健である。 MapLightningは、密なBEVに基づく手法に比べ、中間トークン数を最大で約16.7分の1に減らし、nuScenesとArgoverse 2で最先端の精度と効率を達成する。軽量版はMapTRv2に対して、nuScenesでmAPが10.1、Argoverse 2で16.2上回り、メモリを53%削減しながら推論を1.73倍高速化して40 FPS超を実現する。さらに、不確実性を考慮した地図構築と、後段の軌道予測の改善も示す。コードとモデルは公開予定である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Online vectorized HD map construction is essential for scaling safe autonomous driving and requires accurate, real-time inference. Prior methods typically rely on dense bird's-eye-view (BEV) grids as the intermediate representation. We propose \textit{MapLightning}, which replaces the dense BEV grid with a compact set of 1D learnable map tokens. To construct map tokens from image features, we choose self-attention over vanilla cross-attention because it enables joint interactions and contextual aggregation among image and map tokens. Our transformer-based mapper concatenates map and image tokens, applies full self-attention, discards the image tokens, and retains the updated map tokens for decoding. This design offers three advantages. First, our representation is efficient, using fewer tokens, consuming less memory, and running faster. Second, the lightweight design allows the map decoder to use full rather than deformable cross-attention for better global context. Third, unlike BEV-based methods, our network does not use camera projection parameters, making it robust to camera-extrinsic perturbations. MapLightning uses up to 16.7$\times$ fewer intermediate tokens than dense BEV-based methods and achieves state-of-the-art accuracy and efficiency on nuScenes and Argoverse~2. Its lightweight variant surpasses MapTRv2 by +10.1 mAP on nuScenes and +16.2 mAP on Argoverse~2, while delivering 1.73$\times$ faster inference (40+ FPS) with 53\% less memory. We further show improvements on uncertainty-aware map construction and downstream trajectory prediction. Code and models will be released.
arXiv ID: 2610.01905 / 要約の誤りについて