複数の病理標本をまとめて扱う画像言語モデル
WILSON - a pathology foundation model framework for patient-level analysis and diagnostic text generation
この論文をやさしく読む
ひとことで言うと
患者の複数の病理標本と倍率を一つの画像表現にまとめ、診断文の検索や生成を行うモデルを評価した。
何に役立つ?
考えられる用途は病理症例の情報整理や診断文の検索・下書き支援である。要旨にあるのはモデルの比較評価であり、臨床判断の改善を直接実証したわけではない。
この研究の面白いところ
複数標本を1枚の複数倍率の合成画像として扱い、より大きなモデルに近い結果を小さい計算量で得た。診断文の検索と説明文生成も評価している。
どこまで分かった?
主な比較値は内部コホートで得られた。外部比較は説明文生成の大半について言及されているが、全ての外部条件で優位とは述べていない。患者への効果は要旨からは分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
病理医は倍率を変え、患者1症例の複数の標本を横断して形態を判断する。一方、従来の病理基盤モデルは、単一の標本から多数の小画像を符号化して特徴を集約する。著者らはWILSONを提案する。全標本画像と複数標本からなる症例を、複数倍率を含む1枚の合成画像として表現する画像言語基盤モデルである。Mayo Clinicの約18万9千枚の標本を用い、42臓器・829診断項目にわたる病理報告書を教師情報として学習した。課題別の追加学習なしで、内部の全コホートにおいて症例単位の専用モデルを上回り、macro-F1は0.52対0.38だった。また、最大9.4倍大きい標本単位モデルに匹敵する結果を、272~2,155分の1の計算量で得た。トリプルネガティブ乳がん508症例でモデル全体を追加学習すると、組織学的サブタイプ分類と間質の腫瘍浸潤リンパ球の等級付けが、それぞれmacro-F1で0.16、0.11改善した。診断文の検索ではrecall@1が75.6%で、PRISMの58.1%を上回った。生成した説明文も、内部コホートと外部比較の大半で、PRISMやPRISM2より報告書由来の参照文に近かった。合成画像は病理業務の単位に沿った簡潔な計算表現になり得る。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Pathologists integrate morphology across magnifications and across the slides of a patient case, whereas pathology foundation models encode thousands of tiles from single slides and aggregate their features. Here we present WILSON, a vision--language foundation model that represents whole-slide images and multi-slide cases as single multi-magnification composite images, trained on approximately 189k slides from Mayo Clinic spanning 42 organs and 829 diagnostic entities using pathology reports as supervision. Without task-specific training, WILSON exceeded a dedicated case-level model on all internal cohorts (macro-F1 0.52 versus 0.38) and matched slide-level models up to 9.4 times larger at 272- to 2,155-fold lower compute. End-to-end fine-tuning on 508 triple-negative breast cancer cases improved histologic subtyping and stromal tumor-infiltrating lymphocyte grading by 0.16 and 0.11 macro-F1. WILSON retrieved matching diagnostic text at 75.6% recall@1 (PRISM, 58.1%) and generated captions closer to report-derived references than PRISM and PRISM2 on the internal cohort and on most external comparisons. Composite images thus offer a compact, clinically aligned computational unit for pathology.
著者のコメント
56 pages, 6 main figures, with 11 additional figures and 28 tables in the appendices
arXiv ID: 2609.25123 / 要約の誤りについて