arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

胸部X線でAIの自動判定と医師への委譲を選ぶ方法

Learning to Defer with Guidance on Real World Medical Data

Emma Sun, Joshua Strong and Alison Noble

この論文をやさしく読む

ひとことで言うと

胸部X線の症例ごとにAI判定か人間への委譲か、さらにAIの助言を添えるかを選ぶ方法を評価した。

何に役立つ?

AI読影と臨床医の役割分担を設計・評価する際、助言付き委譲という選択肢の有効性を検討する材料になる。

この研究の面白いところ

固定したAI予測器と学習可能な振り分け器を分け、三つのデータセットと複数の比較条件で評価している。

どこまで分かった?

報告されたのは注釈付きデータセットでの比較であり、臨床現場での安全性や負担軽減を実証したとは要旨に記されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

医用画像の読影は件数が多く時間を要する。AIによる読影は負担を減らし得るが、完全自動での導入には安全性の懸念があり、特異度が低いと実際には臨床医の負担が増える可能性もある。Learning to Defer(L2D)は、入力の特徴とAIモデル・人間の成績を学び、自動予測と人間の専門家への委譲の間で症例を選択的に振り分ける。この方法には理論的保証が示されているが、人間の読影注釈を伴う実世界の医療データでの性能は検証されていなかった。本研究は、AI予測モデルを固定し、学習可能な振り分け・棄却モデルと分離する二段階L2Dの予測器・棄却器形式を、症例ごとに複数人の注釈がある多ラベル胸部X線データセットCollab-CXRで評価する。著者らによれば、人間の注釈を持つ実世界の医用画像データにおけるL2Dを調べた最初の研究である。さらに、判断の選択肢を、自動予測、人間の専門家への委譲、AIの助言を付けた人間の専門家への委譲の三つに広げた、助言付きL2Dを導入する。複数の棄却器の構造、損失関数、利用できる入力特徴の違いを比較し、より大きいVinDr-CXRとCheXpertの二つのデータセットでも再現する。その結果、助言付き二段階L2Dは、従来の二段階L2D、人間のみ、AIのみ、AIの助言を受ける人間の各比較条件を上回った。特に、既存研究で形式的に定義されたL2Dの代理損失関数より単純な損失関数で、この性能を達成した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Medical image interpretation is high-volume and time-consuming, and while AI interpretation can reduce workload, fully autonomous deployment carries potential safety concerns and low specificity may in practice lead to increased clinician workload. Learning to Defer (L2D) addresses this by selectively routing cases between autonomous prediction and human experts by learning from input features and AI model and human performance. While theoretical guarantees have been proven for L2D, its performance has not been validated on real-world medical datasets with human reader annotations. We evaluate the predictor-rejector formulation of two-stage L2D, where the AI predictor model is fixed and separate from the trainable routing or rejector model, on Collab-CXR, a multilabel chest X-ray dataset with multiple human annotations per case. This is the first work to look at L2D in the context of real-world medical imaging data with human annotations. We further introduce a new setup, L2D with Guidance, where the decision space is extended to three choices: predict autonomously, defer to a human expert, or defer to a human expert and provide AI guidance. We compare multiple rejector architectures and loss functions, and different input feature availabilities. This is reproduced on two larger datasets, VinDr-CXR and CheXpert. Our results show that two-stage L2D with Guidance outperforms classic two-stage learning to defer, as well as human-alone, AI-alone and AI-guided human baselines. Notably, this performance is achieved with simpler loss functions compared to formally defined L2D surrogate loss functions in current literature.

著者のコメント

Accepted at HAIC workshop, MICCAI 2026

arXiv ID: 2609.26384 / 要約の誤りについて