arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

放射線画像の報告書でAI評価器Jevが事実の違いを見つける性能

Can Jev Judge Radiology Reports? Evaluating a System One Model for Clinical Factuality

Jiaju Huang, Hao Yang, Xinyu Ma, Xinglong Liang, Kunyan Cai, Junqiang Ma, Shaobin Chen, Yue Sun, and Tao Tan

この論文をやさしく読む

ひとことで言うと

AIが書いた放射線画像の報告書を医師の報告書と比べ、根拠のない記述や省略を検出する評価器を調べた。

何に役立つ?

報告書生成モデルの事実性を評価する際の補助として使える可能性がある。

この研究の面白いところ

各文が相手の報告書に支持されるかを両方向で確認し、質問数と判定費用を抑えた。

どこまで分かった?

二つの専門家データセットなどでの評価であり、臨床的に重要な誤りでは別の評価器がより高い一致を示した。診断の代替を実証したものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

AIが生成した放射線画像の報告書は医師の報告書に似ていても、異常を省いたり、根拠のない所見を加えたり、所見の有無を逆にしたりすることがある。生成器の評価には、こうした事実の違いを測ることが不可欠である。本研究は、System Oneの判断モデルJevを、医師が書いた参照報告書との一致を判定する単純で低費用の評価器として調べる。評価器は各文がもう一方の報告書によって支持されるかを調べ、両方向の判定を組み合わせて、根拠のない主張と省略を捉える。 一つの質問だけを使う設定では、専門家が数えた誤りとのケンドール相関がRadEvalXで0.573、RadEvalExpertで0.398となり、分解と集約の条件をそろえた公開の自然言語推論評価器を上回った。文ごとに一つの支持質問を使う場合は、七つの質問を使う場合と同程度の専門家との一致を保ち、判断に入力するトークンを43~45%減らした。文書化されたAPI価格では、ローカルでの文分解の費用を除き、報告書100組の判定は3セント未満だった。別の制御された誤りの試験では、所見の有無を誤って否定する事例をAUROC 0.977で検出した。 一方、ローカルのRadMatchは、両方の専門家データセットで臨床的に重要な誤りについて、また共通のRadEvalExpert部分集合で全誤りについて、より高い一致を達成した。所見数と誤りの範囲による分析から、ベンチマークとの一致には、医学的誤りの検出だけでなく、報告書の長さと誤りの定義も影響すると分かった。結果は、生成された放射線報告書の事実の違いを測る実用的な判定部品としてJevを支持するとともに、より精緻な評価が有用な場面も示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

An AI-generated radiology report can resemble a physician's report while omitting an abnormality, adding an unsupported finding, or reversing its presence. Measuring these factual differences is essential for evaluating report generators. We study Jev, a System One decision model, as a simple, low-cost judge of agreement with physician-written reference reports. Our evaluator checks whether each statement is supported by the other report and combines these judgments in both directions to capture unsupported claims and omissions. A single-question configuration reaches Kendall correlations of 0.573 on RadEvalX and 0.398 on RadEvalExpert with expert error counts, outperforming an open natural language inference judge under matched decomposition and aggregation. One support question per statement retains similar expert agreement to seven while using 43-45% fewer judgment input tokens. At the documented API price, judgments cost under three cents per hundred report pairs, excluding local decomposition. In a separate controlled-error test, Jev detects false negation with an AUROC of 0.977. Local RadMatch achieves stronger agreement on clinically significant errors in both expert datasets and on total errors in the shared RadEvalExpert subset. Finding-count and error-scope analyses show that benchmark agreement reflects report size and error definitions as well as medical error detection. These results support Jev as a practical judgment component for measuring factual differences in generated radiology reports and identify where more elaborate evaluation remains valuable.

arXiv ID: 2609.27607 / 要約の誤りについて