arXiv論文メモ
新着一覧
cs.CV / cs.LG · 査読状況未確認

複数車両の認識結果を使って環境変化に適応する

Learning from Distributed Eyes: Leveraging Collaborative Perception for Automated Model Adaptation

Yanan Ma, Yihang Tao, Zhengru Fang, Zihan Fang, Yiqin Deng, Xianhao Chen, and Yuguang Fang

この論文をやさしく読む

ひとことで言うと

自分の車だけで作った認識結果を教師にする代わりに、複数車両の情報を合わせた認識を使って、新しい環境にモデルを適応させます。

何に役立つ?

考えられる用途は、人手のラベル付けを増やさずに、環境が変わったときの3次元物体検出を改善することです。

この研究の面白いところ

情報共有の通信量だけでなく、他車には見えても自車には見えない物体や、協調しても残る誤りを個別に扱います。

どこまで分かった?

協調知覚が単独の知覚より良いことが多いという仮定があります。要旨には改善率や帯域幅条件はなく、認識評価の優位性を自動運転全体の安全性実証と同一視できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

自動運転では、環境分布の変化によって認識モデルが新しい環境へ一般化しにくいことがある。教師なしのモデル適応は、手間のかかる人手のラベル付けを不要にする実行可能な解決策だが、自車のデータだけに頼る既存手法は、擬似ラベルの品質が低くなりがちである。この重要な問題に対処するため、協調知覚(CP)をモデル適応のための高品質な教師信号源へ変える、新たな枠組みLDE(Learning from Distributed Eyes)を提案する。 この擬似ラベル付けは、CPが単一エージェントの知覚をしばしば上回るという仮定のもとで、ハイパーパラメータに敏感でなく、比較的信頼できる。しかし単純に実装すると、時間と帯域幅に制限がある中で豊かな特徴を共有する通信上のボトルネック、CPの視野と学習側の視野(FoV)の違い、CPが生成するラベルにも残る不確かさ、という3つの問題が生じる。 これらに対し、適応に最も重要な情報を選択して送る、適応に特化した特徴共有機構、合わないラベルを丁寧に取り除くFoVフィルタリング、擬似ラベルを段階的に活用するカリキュラム学習戦略を設計する。3次元物体検出タスクでの広範な実験は、LDEが事前学習モデルと最先端の教師なし適応法を一貫して上回ることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In autonomous driving, perception models often struggle to generalize to new environments due to domain shifts. While unsupervised model adaptation offers a feasible solution without labor-intensive manual labeling, existing methods that rely solely on the ego-vehicle's data often lead to inferior pseudo-labeling performance. To address this critical issue, we propose LDE, Learning from Distributed ``Eyes", a novel framework that transforms collaborative perception (CP) into a source of high-quality supervision for model adaptation. This pseudo-labeling approach is hyperparameter-insensitive and relatively reliable, assuming CP often outperforms single-agent's perception. However, naively implementing this approach encounters (1) the communication bottleneck of sharing rich features under time and bandwidth constraints, (2) the view discrepancy between the CP view and the learner's Field of View (FoV), and (3) the unreliability even in CP-generated labels. To address these issues, we design an adaptation-oriented feature sharing mechanism that selectively transmits the most critical information for adaptation, an FoV filtering method that meticulously eliminates mismatched labels, and a curriculum learning strategy to progressively exploit pseudo labels. Extensive experiments on 3D object detection tasks demonstrate that LDE consistently outperforms both the pre-trained models and state-of-the-art unsupervised adaptation methods.

著者のコメント

9 pages, 3 figures

arXiv ID: 2609.18511 / 要約の誤りについて