arXiv論文メモ
新着一覧
cs.AI / cs.CV · 査読状況未確認

画像が劣化したときの自動運転向け視覚言語モデルを評価する

Towards Reliable Vision-Language Models for Autonomous Driving

Manasa Mariam Mammen, Priyanka Mary Mammen, Zafer Kayatas and Stefan Wagner

この論文をやさしく読む

ひとことで言うと

画像の品質が落ちたとき、運転場面について答えるAIの正解率と自信の妥当性がどう変わるかを比較しています。

何に役立つ?

自動運転向けVLMを選び評価する際に、画像劣化への強さと、間違ったときの過信を別々に調べるための材料になります。

この研究の面白いところ

5モデルと4データセットに複数の入力形式を組み合わせています。改善法を追加しても、モデルと条件によって効果が異なる点を報告しています。

どこまで分かった?

評価は運転関連の質問応答データセット上のものです。実車での事故率や運転安全性を直接検証したわけではなく、VEAの改善も全条件で共通ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚言語モデル(VLM)は、場面理解、運転に関する推論、意思決定、端から端までの運転など、自動運転のさまざまなタスクで検討されるようになっている。役割が大きくなるほど、頑健性と信頼性の確保が重要になる。実環境では、センサーの不完全さや環境条件によって視覚入力が劣化し、モデルの予測と、それに伴う確信度の両方へ影響する可能性がある。自動運転では、安全に直結する判断のため、正しい予測だけでなく、予測が信頼できない場合を認識することも求められるので、この劣化は特に問題となる。 本研究では、Qwen3.5-9B、Gemma4-E4B、LLaVA-OneVision-7B、DriveFusion/DriveFusionQA-4B、NVIDIA Alpamayo-1.5-10Bの5つのVLMを、運転関連の4つの質問応答データセットで評価する。視覚入力の設定は、単一フレーム、複数視点、複数フレーム、単眼入力を含む。結果は、視覚的な破損の影響がモデル、データセット、入力設定によって異なり、正解率と確信度の信頼性の変化も条件ごとに異なることを示す。 次に、最近提案された推論時の方法Visual Evidence Augmentation(VEA)を適用し、劣化した視覚条件でモデルの信頼性を改善できるかを調べる。VEAは一部のモデルとデータセットで性能を改善するが、すべての設定で一貫した改善が得られるわけではないことが分かった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-Language models (VLMs) are increasingly being explored in autonomous driving for tasks such as scene understanding, driving reasoning, decision-making, and end-to-end driving. As their role becomes more prominent, ensuring their robustness and reliability is increasingly important. In real-world conditions, visual inputs may be degraded by sensor imperfections and environmental conditions, potentially affecting both model predictions and their associated confidence. Such degradation is especially concerning in autonomous driving, where safety-critical decisions require models to make accurate predictions and recognize when their predictions may be unreliable. In this work, we evaluate five VLMs (Qwen3.5-9B, Gemma4-E4B, LLaVA-OneVision-7B, DriveFusion/DriveFusionQA-4B, and NVIDIA Alpamayo-1.5-10B) across four driving-related QA datasets with different visual input settings, including single-frame, multi-view, multi-frame, and monocular inputs. Our results show that the effects of visual corruption vary across models, datasets, and input settings, with changes in accuracy and confidence reliability and also differing across conditions. We then apply Visual Evidence Augmentation ($\mathrm{V}{\scriptstyle \mathrm{EA}}$), a recent inference-time method to examine whether it can improve model reliability under degraded visual conditions. We find that $\mathrm{V}{\scriptstyle \mathrm{EA}}$ improves performance for some models and datasets, although the gains are not consistent across all settings.

arXiv ID: 2610.01531 / 要約の誤りについて