少ない言語注釈で運転指示を学ぶLADA
Less Language, More Latents: Annotation-Efficient VLAs for Driving
この論文をやさしく読む
ひとことで言うと
少量の言語注釈と大量の走行記録を組み合わせ、言語指示に従う運転モデルを学ぶ方法。
何に役立つ?
考えられる用途は、言語注釈を大量に作る負担を抑えた運転モデルの学習である。要旨で実証したのはBench2Driveでの成績である。
この研究の面白いところ
走行の意図を離散的な潜在行動コードにまとめ、少数の言語付き例をそのコードへの橋渡しに使う。
どこまで分かった?
閉ループのBench2Driveで5%未満の言語注釈による成績を示した。実車走行での安全性や性能は要旨に記載がない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚・言語・行動モデル(VLA)は、人の指示に従う自動運転につながるが、学習に必要な自然言語の指示付き映像フレームが少ない。カメラ映像と熟練者の走行軌跡は大量に記録される一方、「交差点で左折」などの言語注釈は少なく、取得にも費用がかかる。この課題に対し、言語注釈のない観測と軌跡の組を言語条件付き制御に利用する3段階の手法LADAを提案する。まずベクトル量子化のボトルネックを持つ潜在行動モデルを学習し、車両の高水準の意図を表す小さなコードブックを作る。次に、言語注釈のある少量のデータで、観測と指示をコードに写す視覚・言語変換器を学習する。最後に、注釈のない全データから観測と潜在行動の組を用いて運転VLAを学習する。言語注釈を5%未満しか使用せず、補助的な思考過程や視覚質問応答データも使わずに、閉ループのBench2Drive評価でDriving Score 87.98、成功率70.46%を達成し、完全教師ありの比較手法と同等以上だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert trajectories are logged at scale, language annotations (e.g., turn left at the intersection) remain scarce and expensive to acquire. To address this challenge, we introduce Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation-trajectory pairs into a substrate for language-conditioned control. First, we train a latent action model with a vector-quantised bottleneck, producing a compact codebook of high-level vehicle intents. Second, a small language-annotated subset is used to train a vision-language translator to map observations and language instructions into this codebook. Third, we train a driving VLA on observation-latent-action pairs over the full unlabelled corpus. Using fewer than 5% of language annotations and without leveraging any auxiliary chain-of-thought reasoning or visual question answering streams, LADA achieves a Driving Score of 87.98 and a Success Rate of 70.46% on the closed-loop Bench2Drive benchmark, matching or surpassing fully supervised baselines.
arXiv ID: 2609.27747 / 要約の誤りについて