arXiv論文メモ
新着一覧
q-bio.QM / cs.AI / cs.CV / cs.LG / eess.IV · 査読状況未確認

複数の病理標本をまとめて扱う画像言語モデル

WILSON - a pathology foundation model framework for patient-level analysis and diagnostic text generation

Saghir Alfasly, Wataru Uegami, Sobhan Hemati, Wenchao Han, Xiaojia Tang, Kevin Thompson, Daniel Stone, Ghazal Alabtah, Saba Yasir, Michael R. Lucas, Eric W. Klee, Cheryl L. Willman, Judy C. Boughey, Matthew P. Goetz, Krishna R. Kalari, H.R. Tizhoosh

この論文をやさしく読む

ひとことで言うと

患者の複数の病理標本と倍率を一つの画像表現にまとめ、診断文の検索や生成を行うモデルを評価した。

何に役立つ?

考えられる用途は病理症例の情報整理や診断文の検索・下書き支援である。要旨にあるのはモデルの比較評価であり、臨床判断の改善を直接実証したわけではない。

この研究の面白いところ

複数標本を1枚の複数倍率の合成画像として扱い、より大きなモデルに近い結果を小さい計算量で得た。診断文の検索と説明文生成も評価している。

どこまで分かった?

主な比較値は内部コホートで得られた。外部比較は説明文生成の大半について言及されているが、全ての外部条件で優位とは述べていない。患者への効果は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

病理医は倍率を変え、患者1症例の複数の標本を横断して形態を判断する。一方、従来の病理基盤モデルは、単一の標本から多数の小画像を符号化して特徴を集約する。著者らはWILSONを提案する。全標本画像と複数標本からなる症例を、複数倍率を含む1枚の合成画像として表現する画像言語基盤モデルである。Mayo Clinicの約18万9千枚の標本を用い、42臓器・829診断項目にわたる病理報告書を教師情報として学習した。課題別の追加学習なしで、内部の全コホートにおいて症例単位の専用モデルを上回り、macro-F1は0.52対0.38だった。また、最大9.4倍大きい標本単位モデルに匹敵する結果を、272~2,155分の1の計算量で得た。トリプルネガティブ乳がん508症例でモデル全体を追加学習すると、組織学的サブタイプ分類と間質の腫瘍浸潤リンパ球の等級付けが、それぞれmacro-F1で0.16、0.11改善した。診断文の検索ではrecall@1が75.6%で、PRISMの58.1%を上回った。生成した説明文も、内部コホートと外部比較の大半で、PRISMやPRISM2より報告書由来の参照文に近かった。合成画像は病理業務の単位に沿った簡潔な計算表現になり得る。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Pathologists integrate morphology across magnifications and across the slides of a patient case, whereas pathology foundation models encode thousands of tiles from single slides and aggregate their features. Here we present WILSON, a vision--language foundation model that represents whole-slide images and multi-slide cases as single multi-magnification composite images, trained on approximately 189k slides from Mayo Clinic spanning 42 organs and 829 diagnostic entities using pathology reports as supervision. Without task-specific training, WILSON exceeded a dedicated case-level model on all internal cohorts (macro-F1 0.52 versus 0.38) and matched slide-level models up to 9.4 times larger at 272- to 2,155-fold lower compute. End-to-end fine-tuning on 508 triple-negative breast cancer cases improved histologic subtyping and stromal tumor-infiltrating lymphocyte grading by 0.16 and 0.11 macro-F1. WILSON retrieved matching diagnostic text at 75.6% recall@1 (PRISM, 58.1%) and generated captions closer to report-derived references than PRISM and PRISM2 on the internal cohort and on most external comparisons. Composite images thus offer a compact, clinically aligned computational unit for pathology.

著者のコメント

56 pages, 6 main figures, with 11 additional figures and 28 tables in the appendices

arXiv ID: 2609.25123 / 要約の誤りについて