arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

診療記録・検査・画像をグラフで結び臨床予測を改善

M2G-LLM: Enhancing Clinical Prediction via Multimodal Graph Reasoning and LLM Context Injection

Inyoung Choi, Sukwon Yun, Jiayi Xin, Jie Peng, Tianlong Chen, Qi Long

この論文をやさしく読む

ひとことで言うと

患者の文章記録だけでなく、検査や画像、受診の時間的なつながりも言語モデルが利用できるようにする方法です。

何に役立つ?

考えられる用途は、複数種類の医療情報を使う予測支援です。評価用データでの予測改善が、患者の転帰改善を直接意味するわけではありません。

この研究の面白いところ

データを単に文章に直すのではなく、患者や受診の関係をグラフで処理し、その情報を言語モデルの中間層に渡しています。

どこまで分かった?

要旨には改善幅、個々の予測課題、施設をまたぐ評価の詳細がありません。報告されているのは二つのデータセットでの比較です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

診療記録、検査結果、医用画像など、多様なデータ形式を統合することは、臨床の意思決定を進歩させるうえで不可欠である。大規模言語モデル(LLM)は非構造化された臨床テキストの処理で優れた性能を示してきたが、テキスト以外の形式を取り込む能力が限られており、医療での幅広い活用を妨げている。 本研究では、グラフニューラルネットワーク(GNN)による複数形式のデータの統合と整合を通じてLLMを強化する、新たな枠組みM2G-LLM(Multimodal MedGraph-LLM)を導入する。この手法は患者の受診間の時間的関係をモデル化し、臨床的に似た患者間で情報を伝播させ、異種のデータ源を整合させることで、情報を充実させたマルチモーダル文脈ベクトルを構築する。これらのベクトルをLLMの中間層に注入し、テキストと非テキストの両形式にわたる共同推論を可能にする。 MIMIC-IVとMIMIC-CXRデータセットでM2G-LLMを評価し、強力なベースラインモデルに比べて臨床予測課題が改善することを示した。結果は、LLMの言語理解とGNNの関係を推論する能力を組み合わせることが、包括的なマルチモーダル医療分析に有望であることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Integrating diverse data modalities --- such as clinical notes, laboratory results, and medical imaging --- is essential for advancing clinical decision-making. While Large Language Models (LLMs) have shown remarkable performance in processing unstructured clinical text, their limited capacity to incorporate non-text modalities hinders their broader utility in healthcare applications. Here, we introduce M2G-LLM (Multimodal MedGraph-LLM), a novel framework that enhances LLMs with multimodal integration and alignment via Graph Neural Networks (GNNs). Our approach models temporal relationships between patient visits, propagates information across clinically similar patients, and aligns heterogeneous data sources to construct enriched multimodal context vectors. These vectors are injected into the intermediate layers of the LLM, enabling joint reasoning over textual and non-textual modalities. We evaluate M2G-LLM on the MIMIC-IV and MIMIC-CXR datasets, demonstrating improvements in clinical prediction tasks over strong baseline models. Our results highlight the promise of combining the language understanding of LLMs with the relational reasoning capabilities of GNNs for comprehensive, multimodal healthcare analysis.

arXiv ID: 2609.21164 / 要約の誤りについて