arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

胸部X線レポートの書き方でAI評価順位が変わる

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny, Nitya M. Bhalla, Fatma Uyar Morency, Pradeep Ravikumar, Zachary C. Lipton, Michael Oberst

この論文をやさしく読む

ひとことで言うと

同じ臨床内容でも参照レポートの表現を変えると、胸部X線レポート生成AIの評価順位が変わることを示しています。

何に役立つ?

レポート生成AIの比較で、臨床内容の正しさと施設ごとの書式への適合を分けて評価するための材料になります。

この研究の面白いところ

正常所見を短くするという記述上の変更だけで、実際のモデル順位が入れ替わる例を示している点です。

どこまで分かった?

比較例はMIMIC-CXR、九モデル、RadCliQ-v1でのものです。参照書換えで評価が変わることは、実際の診断能力が変化したことを意味しません。公開データは120組です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

放射線科医の報告書の書き方は一様ではない。同じ画像を見て同じ臨床所見を特定した二人でも、用語、略記、書式、詳細さの程度によって、表面的には異なる報告書を書くことがある。このような記述慣行の違いは、AIによる放射線画像レポート生成(RRG)モデルの評価で過小評価されている障害である。通常、機械生成レポートは、人が作成した参照レポートとの一致度で評価されるためである。本論文では、既存の評価指標が記述慣行の違いにどれほど敏感かを定量化し、モデル順位を変えるほど大きな影響を明らかにする。 放射線科医の知見を取り入れた報告慣行の変化の分類体系と、臨床的解釈を維持したまま、その分類軸に沿って参照レポートを書き換える手法ReRefを導入する。例えば、MIMIC-CXR上の九つのRRGモデルをRadCliQ-v1で比較すると、参照レポートの正常所見の説明を簡潔にするだけで、Libraは1位から2位へ下がり、CheXOneは3位から1位へ上がる。 この結果は、多くの現行指標が、臨床的解釈と記述慣行への適合を切り分けられていないこと、また、求める記述慣行を正確に反映した「適切な」参照を選ぶことが実務上重要になり得ることを示す。今後の研究を支えるため、MIMIC-CXRから作成した元の参照と別表現の参照の120組からなる、放射線科医が検証したデータセットMIMIC-CXR-Ext-ReRefを公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology, shorthand, formatting, and level of detail. These variations in reporting norms represent an under-appreciated obstacle in efforts to evaluate AI-based radiology report generation (RRG) models, where machine-generated reports are typically assessed based on their concordance with human-generated references. In this paper, we quantify the sensitivity of established evaluation metrics to variations in reporting practices, revealing impacts large enough to alter the rankings of models. We introduce a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of our taxonomy while preserving clinical interpretation. For instance, when comparing the performance of nine RRG models on MIMIC-CXR using RadCliQ-v1, condensing the discussion of normal findings in the reference reports causes Libra to drop from first to second place while CheXOne rises from third to first. Our results suggest that many current metrics fail to decouple clinical interpretation from conformity to reporting practices and that choosing the ``right'' references that accurately reflect the desired reporting practices can be important in practice. To support future research, we release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 (original, alternative) reference report pairs derived from MIMIC-CXR.

著者のコメント

Preprint

arXiv ID: 2609.19093 / 要約の誤りについて