arXiv論文メモ
新着一覧
cs.CV / eess.IV · 査読状況未確認

網膜血管の分割で閾値を選ぶ観察者が評価に与える影響

Observer Choice and Threshold Selection in Retinal Vessel Segmentation: A Subject-Separated Evaluation

Wenhao Xu, Yixian Kong, Ting Pan, Changwei Wang, Feilong Wang, Rongtao Xu

この論文をやさしく読む

ひとことで言うと

網膜血管画像の分割精度が、閾値を選ぶときにどちらの人間の注釈を使うかで変わることを調べた。

何に役立つ?

画像分割の研究で、閾値の選択基準と評価に用いた正解注釈を明記し、モデル間の比較を解釈するために役立つ。

この研究の面白いところ

同じ予測マップやマスクでも、参照する観察者によってDiceが異なった。マキシミン法は多くの適合で閾値を変えたが、低いほうの観察者Diceを改善しなかった。

どこまで分かった?

評価はCHASE DB1の14人・28画像と二人の注釈によるもの。マキシミン法の差の95%区間はゼロをまたぎ、このコホートで精度改善は支持されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

分割の閾値を選ぶ際に使う注釈は評価手順の一部だが、その影響はモデル自体の品質と混同されやすい。本研究は網膜血管分割を対象に、CHASE DB1の全28画像と二人の人間による注釈を使って、この選択を調べる。14人の各被験者の両眼を同じ区分に保つ固定の7分割手順を採用した。観察者1の注釈に対してランダムフォレストとExtra Treesを、3種類の乱数シードで適合させ、合計42回の適合を行った。同一のスコアマップに対し、閾値を0.50に固定する方法、観察者1、観察者2、両観察者の平均に合わせる方法、および画像ごとの低いほうの観察者Diceを最大化するマキシミン法の五つを比較した。ランダムフォレストでは、マキシミン法により21回中19回で閾値が変わったが、低いほうの観察者に対するDiceは70.53%から70.45%へ下がった。対応のある差はマイナス0.073パーセントポイントで、被験者単位の条件付きブートストラップによる95%区間は[マイナス0.384、0.238]だった。Extra Treesでも方向は同じだった。観察者1に合わせて閾値を選んだランダムフォレストの同一マスクは、観察者1に対して73.66%、観察者2に対して71.06%のスコアとなった。結果は、閾値選択時と評価時に参照する注釈を明示して報告することを支持するが、このコホートではマキシミン調整による精度向上を支持しない。分割、元の予測値、指標、コードを公開し、AI支援も開示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The annotation used to select a segmentation threshold is part of the evaluation protocol, yet its effect is easily conflated with model quality. We examine this choice for retinal vessel segmentation using all 28 CHASE DB1 images and both human annotations. A fixed seven-fold protocol keeps both eyes of each of the 14 subjects together. Random forests and Extra Trees are fitted against observer 1 with three random seeds, yielding 42 fits. Five threshold policies share identical score maps: fixed 0.50, observer-1 tuning, observer-2 tuning, mean-observer tuning, and maximin tuning of the per-image lower observer Dice. For random forests, maximin changes the threshold in 19 of 21 fits, but worst-observer Dice decreases from 70.53 percent to 70.45 percent. The paired difference is -0.073 percentage points, with a conditional subject-bootstrap 95 percent interval of [-0.384, 0.238]. Extra Trees shows the same direction. Identical observer-1-tuned random-forest masks score 73.66 percent against observer 1 and 71.06 percent against observer 2. The results support explicit reporting of both the threshold-selection reference and evaluation reference; they do not support an accuracy benefit from maximin tuning in this cohort. All splits, raw predictions, metrics and code are supplied. AI assistance is disclosed.

著者のコメント

7 pages, 3 figures

arXiv ID: 2609.25597 / 要約の誤りについて