物体検出の評価をゆがめるラベルの欠落と曖昧さ
Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty
この論文をやさしく読む
ひとことで言うと
画像の中に本当にある物体に正解ラベルが付いていないと、検出器の正しい出力が誤り扱いされ得ます。そのような評価データの問題を再注釈で調べています。
何に役立つ?
モデルの得点だけを比較する前に、正解データの欠落や人による判断の曖昧さを評価へ取り込むための基盤になります。
この研究の面白いところ
1つの物体に少なくとも11人が関わるソフトラベルを使っています。性能値は変わってもモデルの順位はおおむね安定するという違いも示しています。
どこまで分かった?
最大60%や40%は再注釈による物体数の増加で、検出精度の改善率ではありません。データセット固有の注釈規則も差の一部を説明するとされています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
物体検出は、構造の改善やオープン語彙モデルによって進歩してきたが、本研究は、アノテーションの不完全さがベンチマークの品質を制限している強い証拠を示す。広く用いられる4つのデータセット、COCO、Pascal VOC、Cityscapes、KITTIを再アノテーションすると、ラベル付き物体数が大幅に増加する。例えばKITTIでは最大60%、COCOでは最大40%増え、主因は、以前はラベルが付いていなかった小さい物体、遮蔽された物体、密集した物体である。違いの一部はデータセット固有の注釈規則に由来するが、すべてのデータセットで、ラベルの誤りの主な原因は注釈の欠落であることが一貫して分かる。 高品質なデータを得るため、高い再現率を重視し、物体ごとに少なくとも11人の注釈者から集約したソフトラベルによって曖昧さを捉える、拡張可能なアノテーション工程を導入する。得られた注釈は網羅性を改善し、人間の判断の較正ともよく整合する。モデルの順位はおおむね安定しているものの、ベンチマークの性能値は注釈品質に非常に敏感であることを示す。 不確実性を考慮した物体検出ベンチマークと、実際のラベル誤りに基づくラベル誤り検出ベンチマークという、2つの大規模ベンチマークを導入する。現在の検出器は注釈品質に強く依存し、人間の知覚とずれていることを示す。合成ノイズではよい性能を示すとされてきた現在のラベル誤り検出法も、実際のラベル誤りに対して高い再現率と適合率を得ることに苦戦する。今後の物体検出ベンチマークには、確定的な注釈を超え、有効な物体例を最大限に拾い、現実世界の曖昧さをよりよく反映する、高再現率で不確実性を考慮した評価が必要である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
While object detection has advanced through improved architectures and open-vocabulary models, we provide strong evidence that benchmark quality is limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We show that current detectors are strongly depended on annotation quality and are misaligned with human perception. Current label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation that maximizes valid instances and better reflects real-world ambiguity.
arXiv ID: 2609.21822 / 要約の誤りについて