arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

珍しい運転場面の視覚的根拠と判断を結ぶデータセット

AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

Zhipeng Bao, Wenjie Zhao, Tianle Zhu, Haohua Que, Chence Yang, Geng Yuan, Qianwen Li

この論文をやさしく読む

ひとことで言うと

珍しい運転場面で、判断の根拠となる物体や状況を位置付きで示し、行動理由と軌道計画に結び付けるデータセット。

何に役立つ?

考えられる用途は、自動運転向けVLMの根拠付き推論や軌道予測の学習・評価。要旨は8種類のモデルで指標の改善を報告する。

この研究の面白いところ

重要要素の位置、属性、行動理由、計画を階層的に注釈し、物体サイズを考慮した位置特定指標も導入する。

どこまで分かった?

示された効果はWOD-E2E由来のデータセットと8種類の基盤モデルによる評価。実道路での事故減少や安全性向上を直接示した結果は要旨にない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚言語モデル(VLM)は珍しい自動運転場面への対応に有望だが、既存の運転データセットでは、判断に重要な視覚的根拠を推論や計画に結び付けるための教師情報が限られる。本研究はWOD-E2Eを基に、視覚的根拠に基づく推論データセットAnchorReasoningを導入する。4つの大分類と19の詳細分類にわたる、注釈付きフレーム416,119枚と判断上重要な要素395,379件を含む。各フレームは、重要要素の識別と位置特定、要素の属性と意味、運転行動の理由、行動と軌道の計画を結ぶ、視覚に基づく思考の連鎖(VG-CoT)として整理される。さらに、こうした階層的な能力を順次学ぶカリキュラム型の教師あり微調整法と、物体の大きさを考慮した位置特定品質の評価指標を開発する。 汎用、身体性AI向け、自動運転専用の8種類の基盤モデルで実験した結果、VG-CoTによる教師情報は、視覚的根拠に基づく推論と軌道予測を改善した。モデル全体で、5秒先のADEとFDEはそれぞれ7.84、11.86低下し、RFS FrameとClusterはそれぞれ1.66、1.70向上した。これらの改善は、推論トークンが平均18.5個少なく、推論遅延も1フレーム当たり平均0.32秒短い条件で得られた。珍しい自動運転場面でのVLMの推論と計画に、視覚的根拠と判断に焦点を合わせた教師情報が有用であることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning. We introduce AnchorReasoning, a visually grounded reasoning dataset built on WOD-E2E, containing 416,119 annotated frames and 395,379 decision-critical elements across four major categories and 19 fine-grained types. Each frame is organized as a visually grounded chain-of-thought (VG-CoT) that links decision-critical element identification and localization, element attributes and implications, driving-action rationale, and action and trajectory planning. We further develop a curriculum supervised fine-tuning strategy that progressively learns these hierarchical capabilities, together with an object-size-aware grounding metric for evaluating localization quality. Experiments across eight general-purpose, embodied-AI, and AV-specific backbones show that VG-CoT supervision improves grounded reasoning and trajectory prediction. Across models, 5-s ADE and FDE decrease by 7.84 and 11.86, while RFS Frame and Cluster improve by 1.66 and 1.70. These gains are achieved with 18.5 fewer reasoning tokens and 0.32 s/frame lower inference latency on average, demonstrating the value of visually grounded, decision-focused supervision for VLM reasoning and planning in long-tail autonomous driving.

arXiv ID: 2609.28366 / 要約の誤りについて