arXiv論文メモ
新着一覧
cs.CV / cs.AI / cs.LG · 査読状況未確認

MRIと表形式データを合わせるアルツハイマー病予測

M$^2$PFN: End-to-End Disentangled Alignment for Generalizable Multimodal In-Context Learning in Alzheimer's Disease

Lujia Zhong, Shuo Huang, Jianwei Zhang, Xinyu Nie, and Yonggang Shi

この論文をやさしく読む

ひとことで言うと

脳MRIと表形式の情報を合わせ、アルツハイマー病の分類や認知機能スコアの予測に使う方法です。

何に役立つ?

考えられる用途は、異なる患者集団をまたぐ予測モデルの研究です。要旨で示されたのはADNIと二つの外部集団での予測指標で、診療効果ではありません。

この研究の面白いところ

TabPFNの文脈内学習部分を固定し、MRIと表形式データのエンコーダーを学習して入力の特徴を揃えます。外部集団へ再学習なしで適用しています。

どこまで分かった?

結果はADNIの2240人での三分類と1250人の部分集団でのMMSE回帰、さらにOASIS-3とSCANでの評価です。臨床現場での意思決定や患者転帰の改善は要旨に示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

アルツハイマー病の診断のために画像と表形式データを組み合わせる手法は複数提案されているが、集団をまたぐ一般化には限界がある。TabPFNなどの表形式データ用基盤モデルでは、文脈内学習(ICL)が高い汎化性能と柔軟性を示している。しかしTabPFNは合成した表形式データの事前分布でメタ学習されており、画像から抽出した特徴の統計的な構造とは自然には一致しない。そこで、TabPFNを複数種類のデータを扱うアルツハイマー病予測器に変える、端から端まで学習する枠組みM²PFNを提案する。 M²PFNは、TabPFNのTransformerを通して微分可能な推論を行い、課題の勾配を三次元MRIと表形式データのエンコーダーへ伝える。二つのデータ形式を、情報の分離と対照学習の目的関数を使って、ICLエンジンの事前分布に合う共通の部分空間へ揃える。さらに、表形式データだけによる固定した予測を、学習可能なゲート付きの近道経路として組み込む。ICLエンジンは固定するため、テスト時の一般化に関わる文脈内学習の仕組みは保たれ、端から端までの学習でエンコーダーが利用可能な特徴を形成する。 ADNIの2240人を対象にした正常認知・軽度認知障害・アルツハイマー病の三分類で、マクロF1は65.55%、マクロAUCは82.21%となり、幅広い単一・複数データ形式の比較手法を上回った。出力部だけをTabPFN回帰器へ替えた同じ構成では、1250人の部分集団におけるベースラインのMMSEを予測し、テスト時の平均絶対誤差は1.743で、すべての複数データ形式の比較手法を上回った。再学習をしない二つの外部集団OASIS-3とSCANでも、比較手法の中で最良のAUCと最小のMMSE平均絶対誤差を達成し、認知機能の測定尺度が変わっても移行できた。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

While various multimodal methods combining imaging and tabular data for Alzheimer's disease (AD) diagnosis were proposed, they are often limited in generalization across cohorts. In-context learning (ICL) has demonstrated excellent generalization performances and high flexibility in foundational tabular models such as TabPFN. To extend TabPFN's ICL to multimodal AD analysis, the main obstacle is that TabPFN is meta-trained on synthetic tabular priors that do not naturally match the statistical structure of image-derived features. We propose M$^2$PFN, an end-to-end framework that turns this tabular foundation model into a multimodal AD predictor. M$^2$PFN (i) performs differentiable inference through TabPFN's transformer, back-propagating task gradients into 3D-MRI and tabular encoders; (ii) aligns the two modalities into a shared subspace, via disentanglement and a contrastive objective, matched to the ICL engine's prior; and (iii) folds in a frozen tabular-only prediction through a learnable gated shortcut. Because the ICL engine stays frozen, its in-context mechanism is preserved for test-time generalization, while end-to-end training shapes the encoders into features it can exploit. On ADNI ($n=2240$, three-class CN/MCI/AD), M$^2$PFN attains $65.55\%$ macro-F1 and $82.21\%$ macro-AUC, surpassing a comprehensive set of unimodal and multimodal baselines. By swapping only the head for a TabPFN regressor, the same architecture regresses baseline MMSE on a $1250$-subject sub-cohort to test MAE $1.743$, outperforming every multimodal baseline. On two external cohorts (OASIS-3 and SCAN) with no retraining, M$^2$PFN achieves the best AUC and the lowest MMSE MAE across all baselines, and transfers even when the cognitive instrument changes.

著者のコメント

Under review

arXiv ID: 2609.28836 / 要約の誤りについて