画像や表の欠測に対応してAlzheimer病の経過を予測
MMAP: Multimodal Missing-Aware Pretraining for Longitudinal Alzheimer's Prediction
この論文をやさしく読む
ひとことで言うと
画像検査や表形式の診療データに欠けがあっても、それらを組み合わせてAlzheimer病の経過やアミロイド状態を予測する方法です。
何に役立つ?
考えられる用途は、すべての検査がそろわない医療データで進行予測モデルを学習・利用することです。研究では病期移行とアミロイド状態の予測課題を評価しています。
この研究の面白いところ
画像と表で異なる事前学習を使い、欠測を表すトークン生成器で双方をつなぎます。疾患ラベルだけに頼らず、対照学習と再構成も利用しています。
どこまで分かった?
要旨には患者数、欠測率、改善幅、外部施設での評価は示されていません。予測課題で基準手法を上回ったという結果で、治療効果や臨床運用上の利益を実証したものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
臨床的な意思決定は、複数種類の医療データで特徴付けられる患者の健康状態を理解し、疾患の進行経路を予測することに大きく依存する。AIには、複数モダリティの医療データから有用な表現を学び、病状の進行予測と臨床判断を助ける大きな可能性がある。しかし、医療データによくあるモダリティの欠落や表形式データの不完全さが、予測AIモデルの開発を制約する。また、疾患ラベルだけでは、高次元の複数モダリティから表現を学ぶための教師信号が限られることもある。 本研究では、不完全なデータから画像と表の表現を学ぶ、新たな複数モダリティ欠測対応整合事前学習法MMAPを提案する。画像エンコーダーは、効率的なシグモイド対照学習と生成的な再構成を組み合わせて事前学習する。表形式エンコーダーは表形式基盤モデルの上に構築する。欠測トークン生成器により、両エンコーダーが不完全なデータを入力として受け取れるようにし、画像が欠けた場合と表形式データが欠けた場合のいずれにもモデルが頑健であるようにする。学んだ複数モダリティ表現の臨床的な有用性を、Alzheimer病の2つの難しい縦断的課題、病期の移行予測とアミロイド状態の予測で評価する。提案手法は、強力な複数モダリティおよび単一モダリティの基準手法を上回る。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Clinical decision making heavily relies on predicting the disease progression trajectory by seeking to understand patient's health status which is characterised by multimodal medical data. AI holds great potential for learning useful representations from multimodal medical data to predict disease progression and aid clinical decision making. However, development of predictive AI models is constrained by missing modalities and incomplete tabular data frequently occurring in medical datasets. In addition, disease labels alone may only provide limited supervisory signals for learning representations from high-dimensional multimodal data. Here, we present MMAP, a novel Multimodal Missing-aware Alignment Pretraining method for learning image-tabular representations from incomplete data. An image encoder is pretrained with efficient sigmoid contrastive learning combined with generative reconstruction. A tabular encoder is built upon a tabular foundation model. A missing token generator enables the two encoders to take incomplete data as input, enabling the model to be robust against missing modalities, either with missing images or missing tabular data. We evaluate the clinical usefulness of the learnt multimodal representations on two challenging longitudinal clinical tasks for Alzheimer's disease: predicting disease stage conversion and predicting amyloid status. The proposed method outperforms strong multimodal and unimodal baselines.
著者のコメント
To be published in the proceedings of the 2026 MICCAI Workshop on Multimodal Learning with Medical Tabular Data
arXiv ID: 2609.26617 / 要約の誤りについて