arXiv論文メモ
新着一覧
cs.CV / cs.SE · 査読状況未確認

自動車の物体認識に使う視覚言語モデルの事前評価

Towards Systematic Qualification of Vision-Language Models for Automotive Perception Systems

Malsha Ashani Mahawatta Dona, Konstantinos Rokanas, Alexander Säfström, Krishna Ronanki, Christian Berger

この論文をやさしく読む

ひとことで言うと

車載認識に使う視覚言語モデルが物体を誤って報告する頻度を、設計段階で比較できる評価手順です。

何に役立つ?

車載モデルの比較や採用判断の材料になります。研究で実証したのは、nuScenes上の3モデルについて幻覚を再現可能に数えることです。

この研究の面白いところ

安全に関係する物体の注釈を固定した分類体系で揃え、同義語も考慮して評価します。実行時の監視を設計時の評価で補う構成です。

どこまで分かった?

要旨の評価対象はnuScenesデータセット上の3種類のVLMです。実車での安全性改善や事故削減は要旨で示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

人工知能は多くの分野で採用され、近年進歩した視覚言語モデル(VLM)も、車両の周囲認識や安全性保証を支援する方法として検討されている。しかし、こうしたモデルは幻覚を起こし得るため、組み込まれた自動車システムの安全性を脅かす可能性がある。自動車分野では、存在しない交通物体を報告するだけでなく、実際にある物体を見落とすことも危険につながり得る。安全で信頼できるAIの検証・妥当性確認手法の研究は増えているものの、多くは実行時か設計時の一方だけを扱っており、安全性が重要な実際の車両認識には不十分かもしれない。本論文はHuangらの分類法に基づいて設計時と実行時の検証・妥当性確認技術を分析し、実行時監視を補完する設計時の適格性評価手順を提案する自動車分野の研究を示す。手順では、安全上重要な対象について固定されたオントロジーに基づく構造化注釈と、同義語を考慮した評価を組み合わせ、nuScenesデータセット上で最先端のVLM3種類を統計的に評価する。提案手法によって、自動車の周囲認識タスクでVLMが生じさせる幻覚を、決定的かつ反復可能な形で定量化できることを観察した。この手順は、設計時の検証・妥当性確認におけるモデル比較や導入を見据えた技術判断を支え、信頼できる自動車認識に向けた包括的な検証戦略に寄与する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The field of Artificial Intelligence has been adopted for many application domains. Vision Language Models are one of the recently advanced AI techniques that have been explored to support automotive features such as vehicle perception, and safety assurance. However, such language models are prone to hallucinations, posing a potential threat to the safety of automotive systems that may incorporate them. Within the automotive domain, VLMs could not only hallucinate traffic objects, but could also fail to identify traffic objects that are actually present, which may potentially lead to dangerous situations. Though we have observed a growing body of literature that proposes verification and validation techniques for safe and trustworthy AI, these methods are often studied in isolation, focusing either on run-time or design-time phases. Such isolated techniques could be insufficient in safety-critical, realistic contexts such as automotive perception systems. In this paper, we analyze design-time and run-time verification and validation techniques based on a taxonomy presented by Huang et al. We present an automotive study in which a design-time qualification workflow is proposed to complement run-time monitoring. This workflow combines a fixed safety-relevant ontology-based structured annotation system together with a synonym-based evaluation process to statistically evaluate three state-of-the-art VLMs against data from the nuScenes dataset. We observed that the proposed technique enables deterministic and repeatable quantification of the hallucinations VLMs generate in automotive perception-related tasks. The proposed workflow supports model comparison and deployment-oriented engineering decisions within the design-time verification and validation process and will contribute to a holistic verification strategy that strives towards trustworthy automotive perception systems

著者のコメント

Accepted in ICTSS 2026 - 38th International Conference on Testing Software and Systems

arXiv ID: 2609.25945 / 要約の誤りについて