arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

基盤モデルの確信度と単眼深度で未知物体を領域検出

Combining Foundation Model Confidence and Monocular Depth for Training-Free Out-of-Distribution Segmentation

Serin Varghese, Fabian Hüger, Kira Maag

この論文をやさしく読む

ひとことで言うと

自動運転の画像で、学習時に見ていない物体を、モデルの確信度と深度情報から追加学習なしで見つける研究。

何に役立つ?

未知の物体を検出する車載知覚システムの評価や、追加学習を抑えた方法の検討に役立つ。

この研究の面白いところ

領域分割モデルの確信度に、単眼深度から得る幾何情報を組み合わせ、画素ごとの未知物体スコアを作る。

どこまで分かった?

要旨は道路中心のベンチマークで良好と述べるが、具体的な精度値や比較対象は記していない。実車での安全性を実証したものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

自動運転車が開かれた環境で走ると、珍しい動物や落下した荷物など、事前に知られていない物体に直面する。こうした学習時の分布から外れた物体を確実に検出し、その領域を分けることは、周囲の安全な理解と判断に重要である。既存手法の多くは、分布外データの学習例、領域分割モデルの再学習、または専用の補助構造を必要とし、実用上の制約となる。 本研究は、特定の課題向けの微調整や異常データを使わず、基盤となる領域分割モデルの確信度予測から、画素ごとの分布外スコアを直接求める、追加学習不要の方法を提案する。分布外の領域分割を頑健にするため、単眼画像から推定した深度の幾何学的情報も判断に組み込み、不確実性による予測を補完する。SegmentMeIfYouCanベンチマークで評価し、現実の知覚システムの時間的な性質を反映するため、動画での分布外物体の追跡性能も調べた。道路を中心としたベンチマークで良好な性能を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Autonomous vehicles operating in open-world scenarios are inevitably confronted with previously unknown objects, such as exotic animals or loose cargo. The reliable detection and segmentation of these out-of-distribution (OOD) objects is therefore crucial for a safe understanding of the environment and decision-making. Most existing approaches require access to OOD training samples, retraining of the segmentation backbone, or dedicated auxiliary architectures, limiting their practical applicability. We propose a training-free method that derives dense OOD scores directly from the confidence predictions of a foundation segmentation model, without any task-specific fine-tuning or access to anomalous data. To improve the robustness of our OOD segmentation, geometric information from monocular depth estimation is incorporated into the decision process, providing complementary cues to uncertainty-based predictions. We evaluate the proposed method on the SegmentMeIfYouCan benchmark and additionally assess its performance on OOD tracking in video sequences, reflecting the temporal nature of real-world perception systems. The method performs strongly on road-centered benchmarks.

arXiv ID: 2609.22896 / 要約の誤りについて