arXiv論文メモ
新着一覧
eess.IV / cs.CV · 査読状況未確認

骨腫瘍X線画像の局所像と全体像から領域と種類を推定

Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis

S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury, Adam Mushtak, Israa Al-Hashimi, Sohaib Bassam Zoghoul

この論文をやさしく読む

ひとことで言うと

骨腫瘍のX線画像で、病変の拡大像と骨全体の画像を組み合わせて領域と種類を推定した研究。

何に役立つ?

画像診断の支援モデルを設計する際、局所の細部と全体の位置関係を併用する方法の評価材料になる。

この研究の面白いところ

3,746件の複数施設データで患者単位に試験データを分け、領域分割と分類を同時に評価した。

どこまで分かった?

報告値は保持した試験データでのモデル性能であり、実際の診療で診断の早さや患者の結果が改善したことは要旨で示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

原発性骨腫瘍はまれだが臨床的に進行性の腫瘍であり、X線画像からの診断は形態の多様さ、病変の境界の見えにくさ、骨構造の重なりのために難しい。単一の画像範囲だけを使う既存モデルの限界に対応するため、病変を切り出した画像とX線画像全体の間で双方向のクロスモーダル注意機構を用い、病変領域の分割と腫瘍の種類の分類を同時に行う、二入力・複数課題の学習枠組みを提示する。著者らの知る限り、この組合せは初めてである。複数施設から集めたBone Tumor X-ray Radiograph Dataset(BTXRD、3,746件)を使い、YOLOに基づく検出器で関心領域を作り、元画像全体と組にして二経路のDenseNet121に入力する。新しいクロスモーダル注意融合で特徴を統合し、階層的・多尺度の特徴融合でさらに精緻化して、病変の細部と解剖学的な全体の文脈を両立させる。 患者単位で分離して保持した試験データで評価すると、単一入力の基準モデルより高い性能を示し、全体のDice類似係数0.896、分類のマクロ平均F1スコア0.928を達成した。特に悪性の骨肉腫に対する感度が高く、AUCは0.999だった。これらの結果は、局所像と全体像を使う方法が、放射線科医の正確で早い判断を支援する可能性を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Primary bone tumors are rare but clinically aggressive neoplasms whose diagnosis from radiographs is challenged by heterogeneous morphology, subtle lesion margins, and overlapping bone structures. To address the limitations of existing single-view models, we present a dual-input, multi-task learning framework that, to our knowledge, is the first to apply bidirectional cross-modal attention between a lesion crop and the full radiograph for joint segmentation and subtype classification. Using the multi-institutional Bone Tumor X-ray Radiograph Dataset (BTXRD, n=3,746), we employ a YOLO-based detector to generate regions of interest, which are paired with full images as inputs to a dual-stream DenseNet121 architecture. Features are integrated via a novel cross-modal attention fusion strategy, refined by Hierarchical Multi-scale Feature Fusion, effectively balancing fine-grained lesion detail with global anatomical context. Evaluated on a held-out patient-level test split, the model demonstrates superior performance over single-input baselines, achieving an overall Dice Similarity Coefficient of 0.896 and a macro-averaged classification F1-score of 0.928. Notably, the system exhibits exceptional sensitivity for malignant osteosarcoma (AUC 0.999), validating the potential of dual-stream context modeling to support radiologists in accurate, early decision-making.

arXiv ID: 2609.28732 / 要約の誤りについて