arXiv論文メモ
新着一覧
cs.CV / eess.IV · 査読状況未確認

三次元放射線画像向けの畳み込み型・Transformer型基盤モデル

nnFoundation: 3D Foundation Models for Radiology

Constantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, Bálint Kovács, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica Kächele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Krämer, Maximilian Zenk, Tim Rädsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David Kügler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina Decker, Mohammad Yaqub, Arnd Heuser, Martin Reuter, Julia A. Schnabel, Tobias Heimann, Florin Ghesu, Paul Brachmann, Claus P. Heußel, Alexander Radbruch, Gianluca Brugnara, Aditya Rastogi, Martha Foltyn-Dumitru, Heinz-Peter Schlemmer, Ignaz Reicht, Julius C. Holzschuh, Michael Bach, Bram Stieltjes, Kai Schlamp, Lena Maier-Hein, Marco Nolden, Ralf Floca, Paul F. Jäger, Philipp Vollmuth, Fabian Isensee, Klaus H. Maier-Hein

この論文をやさしく読む

ひとことで言うと

CT・MRI・PETの210万件で学習した二種類の三次元基盤モデルを、108の画像課題で比較した。

何に役立つ?

放射線画像の課題に応じて事前学習モデルの種類を選ぶ際の参考になる。

この研究の面白いところ

局所的な画像課題では畳み込み型、大域的な意味理解ではTransformer型が優れるという違いを示した。

どこまで分かった?

多数の課題で評価されたが、要旨は臨床現場での患者への効果を検証したとは述べていない。課題とデータに応じた適応が必要である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

放射線画像を扱う人工知能は急速に進歩しているが、多くのシステムは用途が狭く、大量のデータを必要とし、データの分布が変わると弱い。基盤モデルは、より転用しやすく少ないデータで使える解決策となり得るが、従来の方法は規模が限られ、評価範囲も狭く、一つの事前学習モデルで多様な後続課題に対応できると仮定することが多い。本研究は、相補的な畳み込み型とTransformer型の三次元放射線画像基盤モデルnnFoundationを提示する。Human Radiome Project(THRP)の中で開発され、125の施設・公開データセットから集めた210万件のCT、MRI、PET画像ボリュームで学習した。 領域分割、検出、分類、報告書生成、画像検索にわたる108課題で評価し、分布が変わる条件、外部協力者による評価、データや計算資源が少ない条件も含めた。畳み込み型とTransformer型のnnFoundationは、すべての課題種別において、従来の三次元基盤モデルとゼロからの学習を一貫して上回り、放射線画像で最先端の性能に達した。ただし、性能は課題に依存する一貫した構造を持つ。畳み込み型は空間的に局所化した課題で優れ、Transformer型は大域的な意味理解を要する課題と、特徴を固定して使う設定で優れた。事前学習後に、データセットの性質に合わせて基盤モデルの構造を動的に調整すると、異なる三次元データへの転用がさらに改善した。転用可能な三次元放射線画像の性能は、単一の万能モデルではなく、規模を拡張できる事前学習、相補的な構造、データに合わせた適応の組み合わせによって決まることを示す。nnFoundationのモデルをnnU-NetとnnDetectionに組み込んで公開し、既存の放射線画像ワークフローですぐ使えるようにした。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Radiological artificial intelligence has advanced rapidly, yet most systems remain narrowly task-specific, data-intensive, and fragile under domain shift. Foundation models promise more transferable and data-efficient solutions, but existing approaches are limited in scale, evaluated narrowly, and often assume that a single pretrained model can support diverse downstream tasks. Here we present nnFoundation, complementary convolutional and transformer-based 3D radiological foundation models. Developed within the Human Radiome Project (THRP), nnFoundation is trained on 2.1 million CT, MRI, and PET image volumes from 125 institutional and public datasets. We evaluate them across 108 tasks spanning segmentation, detection, classification, report generation, and image retrieval, including evaluations under domain shift, by external partners and in low-data and low-compute regimes. Across all task types, our convolution- and transformer-based nnFoundation models consistently outperform both prior 3D foundation models and training from scratch, establishing state-of-the-art performance for radiological imaging. However, performance follows a consistent task-dependent structure: the convolutional nnFoundation model dominates spatially localized tasks, whereas the transformer-based nnFoundation model excels in tasks requiring global semantic reasoning and in frozen-feature settings. Dynamically aligning the foundation model topology with the dataset characteristics post-hoc further improves transfer across heterogeneous 3D settings. These results show that transferable 3D radiological performance is governed not by a single universal model, but by the interplay of scalable pretraining, complementary architectures, and dataset-aware adaptation. We release nnFoundation models integrated into nnU-Net and nnDetection, enabling immediate application across established radiology workflows.

arXiv ID: 2609.26924 / 要約の誤りについて