arXiv論文メモ
新着一覧
cs.CV / cs.RO · 査読状況未確認

画像基盤モデルで雑音に強い3次元物体検出を学習

Towards robust multimodal 3D object detection via visual foundation models

Ziying Song, Lin Liu, Hongyu Pan, Shaoqing Xu, Lei Yang, Mingzhe Guo, Caiyan Jia

この論文をやさしく読む

ひとことで言うと

画像モデルが持つ知識を、カメラとLiDARを使う物体検出へ取り込みます。特徴の多段階処理、雑音抑制、軽量モデルへの知識蒸留を組み合わせた方法です。

何に役立つ?

考えられる用途は、天候やセンサー雑音が変わる自動運転環境の物体検出です。要旨で報告している実証は27種類の分布外劣化条件での検出評価であり、自動運転車全体の安全性評価ではありません。

この研究の面白いところ

SAMをそのまま接続するのではなく、自動運転向けの微調整と特徴融合を行ったうえで、点群ネットワークへ知識を移します。高周波雑音の抑制にも専用の機構を設けています。

どこまで分かった?

性能は全条件で一律に優れるとは述べられず、概して優れるか同等に競争力があるとの報告です。要旨にはデータセット別の数値、処理速度、実車走行での評価結果は記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

複数種類のセンサーを用いる3次元物体検出は、LiDARとカメラの相補的な情報を統合するため、自動運転における頑健な認識の基盤となる。しかし既存手法は、センサー雑音、悪天候、環境変化によって生じる分布外(OOD)の劣化に対して、頑健性を維持できないことが多い。この問題に対し、Segment Anything Model(SAM)などの視覚基盤モデル(VFM)を活用する、頑健性と汎化能力を備えたマルチモーダル3次元物体検出フレームワークRoboDistillを提案する。 第一に、SAMを自動運転画像で微調整し、意味情報の豊かな特徴表現を抽出する、分野特化の事前学習戦略SAM-ADを導入する。第二に、SAMの特徴を複数のスケールで精緻化・アップサンプリングし、LiDAR特徴と円滑に融合するAD Feature Pyramid Network(AD-FPN)を設計する。第三に、重要な文脈情報を保ちながら高周波のセンサー雑音を抑制するDepth-Guided Wavelet Attention(DGWA)モジュールを開発する。最後に、事前学習済みのSAM-ADを教師とし、高品質な視覚知識を軽量な点群ネットワークへ蒸留するKD Fusionを導入し、雑音下での頑健性を高める。 27種類の厳しいOOD劣化条件で行った広範な実験では、RoboDistillは代表的な最先端手法に対し、概して上回るか競争力のある検出性能と頑健性を示した。本研究はVFMと3次元物体検出の間をつなぎ、実世界の自動運転への応用に向けた頑健なマルチモーダル認識を前進させる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Multimodal 3D object detection is fundamental to robust perception in autonomous driving because it integrates complementary information from LiDAR and camera sensors. However, existing methods often fail to maintain robustness under out-of-distribution (OOD) corruptions caused by sensor noise, adverse weather, and environmental changes. To address this problem, we propose RoboDistill, a robust and generalizable multimodal 3D object detection framework that leverages visual foundation models (VFMs), such as the Segment Anything Model (SAM). First, we introduce SAM-AD, a domain-specific pretraining strategy that fine-tunes SAM on autonomous-driving imagery to extract feature representations with rich semantic information. Second, we design the AD Feature Pyramid Network (AD-FPN) to refine and upsample SAM features at multiple scales for seamless fusion with LiDAR features. Third, we develop the Depth-Guided Wavelet Attention (DGWA) module, which suppresses high-frequency sensor noise while preserving critical contextual information. Finally, we introduce KD Fusion, in which the pretrained SAM-AD serves as a teacher that distills high-quality visual knowledge into a lightweight point-cloud network, thereby improving robustness under noisy conditions. Extensive experiments across 27 challenging OOD corruption settings show that RoboDistill generally delivers stronger or competitive detection performance and robustness relative to representative state-of-the-art methods. This work bridges the gap between VFMs and 3D object detection and advances robust multimodal perception for real-world autonomous-driving applications.

著者のコメント

27 pages, 5 figures. Bilingual English-Chinese manuscript; the complete English version appears first, followed by the complete Chinese version

arXiv ID: 2609.23541 / 要約の誤りについて