arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

胸腹部CTの3次元画像と言語を結ぶ解析モデル

NV-Reason-CT: 3D Visual Language Model for CT Analysis

Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu

この論文をやさしく読む

ひとことで言うと

胸腹部CTを3次元のまま読み取り、所見や報告書を言語で出す研究用モデルを開発・評価した。

何に役立つ?

考えられる用途は、放射線科医がCT所見を確認する際の研究支援である。要旨には分類・報告書生成の評価と、専門家による予備的な時間評価が示される。

この研究の面白いところ

画像を平面ごとにまとめず、全ての視覚トークンと3次元座標を言語生成へ渡す構成を採る。

どこまで分かった?

読影時間の50%短縮は予備的研究で報告された結果である。要旨は臨床での安全性や診断への導入を証明していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

本研究は、胸部と腹部のCT向けに、3次元画像をそのまま符号化し、放射線科医の指導に基づく推論を組み合わせた生成型の視覚言語モデルNV-Reason-CTを提示する。3次元の視覚トランスフォーマーと言語モデルを結合し、全ての視覚トークンと明示的な3次元座標を、空間的なトークンをさらに統合せずに言語生成へ渡す。これにより、画像符号化器の内部と、テキストと共同処理する際の言語モデルの位置符号化の双方で、体積画像の空間情報を保つ。学習には、70,111件の異なるCT画像入力に由来する約55万件の複数形式の指示例を整理したデータ群を用いる。標準化した報告書、異常と解剖部位に焦点を当てた質問、複数ターンの対話、専門家のCT読影を記録・書き起こして得た放射線科医による推論を組み合わせた。専門家の注釈は直接の教師信号となり、報告書に基づく追加の合成推論も導く。端から端までの教師あり微調整の後、胸腹部の異常集合に関する検証可能な報酬を用いたGroup Relative Policy Optimizationを行う。 モデルは異常の分類、報告書の生成、確認可能な所見、鑑別診断、不確実性を伴う対話的な推論に対応する。公開CTベンチマークと、学習から除いたNIHの集団で評価した。CT-RATEでは、課題専用の分類ヘッドなしでマクロF1が0.614、マクロAUROCが0.871となり、生成報告書から算出したマクロF1は0.592だった。放射線科医による予備的研究では、AIを使った確認に好意的な確信度評価が得られ、報告された読影と報告書作成の平均時間は50%短かった。体積医用画像向けの説明可能なAI研究の再現を支援するため、モデルと学習コードを公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We present NV-Reason-CT, a generative vision--language model for chest and abdominal CT combining native 3D visual encoding with radiologist-guided reasoning. The model couples a native 3D vision transformer with a language model, passing all visual tokens and their explicit 3D coordinates into language decoding without further spatial token merging. This retains volumetric spatial information within the vision encoder and through the language model's positional encoding during joint processing with text. We train on a curated corpus of approximately 550,000 multimodal instruction examples from 70,111 unique CT image inputs, combining standardized reports, abnormality-focused and anatomy-specific questions, multi-turn interactions, and radiologist-authored reasoning from recorded and transcribed expert CT interpretations. Expert annotations provide direct supervision and guide additional report-grounded synthetic reasoning. End-to-end supervised fine-tuning (SFT) is followed by Group Relative Policy Optimization (GRPO), with verifiable rewards over chest and abdominal abnormality sets. The model supports abnormality classification, report generation, and interactive reasoning with reviewable observations, differential diagnoses, and uncertainty. Evaluation spans public CT benchmarks and a held-out NIH cohort. On CT-RATE, NV-Reason-CT achieves a macro-F1 of 0.614 and macro-AUROC of 0.871 without a task-specific classification head; generated reports achieve a report-derived macro-F1 of 0.592. In a preliminary study with expert radiologists, AI-assisted review received favorable confidence ratings and was associated with a 50% reduction in average reported interpretation and reporting time. We release the model and training code to support reproducible research on explainable AI for volumetric medical imaging.

arXiv ID: 2609.27511 / 要約の誤りについて