arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

まれな胸部X線所見を識別する画像言語モデル

Med-AR: Autoregressive Vision-Language Pretraining for Long-Tailed Chest X-Ray Classification and Uncertainty-Aware Evaluation

Janhavi Prabhu, Sahil, Akshay V, Shivam Shukla, Manoj Tadepalli, and Preetham Putha

この論文をやさしく読む

ひとことで言うと

胸部X線画像で、よくある異常だけでなく、まれな所見も識別するための事前学習モデルです。

何に役立つ?

胸部X線分類モデルの研究で、頻度の低いラベルや判定を保留する能力も含めて比較する際に役立ちます。臨床での診断性能を実証した結果ではありません。

この研究の面白いところ

公開データではまれな所見の指標が改善しましたが、内部データでは比較手法が優れる指標もあり、評価条件による違いを示しています。

どこまで分かった?

公開データセット上の改善は内部データで一様には再現されていません。要旨の結果は分類評価であり、臨床利用の妥当性を示すものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

出現頻度に偏りがある胸部X線画像の分類には、よくある異常と、まれで微妙な所見の両方を捉える視覚表現が必要である。そこで、構造化された読影報告、異常に焦点を当てた文章、領域注釈で事前学習した、放射線画像に特化した自己回帰型の視覚言語モデルMed-AR-8BとMed-AR-2Bを提案する。共通のML-Decoder分類ヘッドを使い、画像エンコーダーを複数ラベル分類に移したときの性能を、Med-CLIP、CheXFound、EVA-Base、ARK、BioViL-Tを含む対照学習、自己教師あり学習、教師あり学習の事前学習済みエンコーダーと比較した。細かな認識能力を調べるため、MIMIC-CXRとCheXpertについて、報告から作り、大規模言語モデルで拡張したラベル集合も構築した。PadChest、MIMIC-CXR、CheXpertでは、Med-AR-8Bは頻出、中程度、まれな所見の各群で、平均AUROCと平均AUPRCの両方がMed-CLIPを上回った。MIMIC-CXRのまれなラベル群では、平均AUPRCが0.1033から0.1441に上がった。Med-AR-2BはPadChestで最も高い識別結果を得た。広い範囲のエンコーダー比較でも、各公開データセットの報告されたすべての頻度群で、Med-ARのいずれかが平均AUROCと平均AUPRCの最高値を得た。両モデルは三つの公開データセットすべてで、Med-CLIPよりリスク・カバレッジ曲線の余剰面積が小さく、評価した手順のもとで選択的予測の性能が改善したことを示す。一方、内部データでの結果は指標によって異なり、全体およびまれなラベルのAUPRCと選択的予測ではMed-CLIPが優位だった。これらの知見は、評価した公開ベンチマークでMed-ARが偏りのある胸部X線分類の有力な事前学習方法であり、識別性能と選択的予測をともに評価する価値があることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Long-tailed chest X-ray classification requires visual representations that capture both common abnormalities and subtle, infrequent findings. We propose Med-AR-8B and Med-AR-2B, two radiology-native autoregressive vision-language models pretrained with structured reports, abnormality-focused text, and region annotations. We evaluate the transfer of their visual encoders to multi-label classification against contrastive, self-supervised, and supervised pretrained encoders, including Med-CLIP, CheXFound, EVA-Base, ARK, and BioViL-T, using a common ML-Decoder classification head. To assess fine-grained recognition, we also construct LLM-expanded, report-derived label sets for MIMIC-CXR and CheXpert. Across PadChest, MIMIC-CXR, and CheXpert, Med-AR-8B outperforms Med-CLIP in mean AUROC and AUPRC for head, medium, and tail findings. On MIMIC-CXR, it increases tail-label mean AUPRC from 0.1033 to 0.1441. Med-AR-2B achieves the strongest discrimination results on PadChest. Across the broader encoder comparison, a Med-AR variant achieves the highest mean AUROC and AUPRC in every reported prevalence group on each public dataset. Both Med-AR variants also achieve lower excess area under the risk-coverage curve than Med-CLIP on all three public datasets, indicating improved selective-prediction performance under the evaluated protocol. Internal results are metric-dependent, with Med-CLIP retaining advantages in overall and tail AUPRC and in selective prediction. These findings establish Med-AR as a strong pretraining recipe for long-tailed chest X-ray classification on the evaluated public benchmarks and demonstrate the value of assessing discrimination and selective prediction together.

著者のコメント

80 pages including supplementary material, 28 figures, and 22 tables. Supplementary material is included

arXiv ID: 2609.29156 / 要約の誤りについて