arXiv論文メモ
新着一覧
eess.IV / cs.CV · 査読状況未確認

GAN生成画像の追加は脳腫瘍MRI分類を改善するか

Does DCGAN-Based Synthetic Augmentation Improve Brain Tumor MRI Classification? An Empirical Study

Irhum Jawad Khan, Talha bin Aslam

この論文をやさしく読む

ひとことで言うと

脳MRIの分類にGANで作った画像を加えても、今回の条件では正解率が上がらなかったという比較研究。

何に役立つ?

医用画像の学習データを合成画像で増やす際に、見た目や分布の近さだけでなく、同じ試験データで分類性能も確認する必要があると判断する材料になる。

この研究の面白いところ

4クラス計7200枚のMRIを使い、分類器と試験データを固定して、実画像のみと各クラス500枚の合成画像を追加した条件を比べた。正解率はともに96%で、ROC-AUCは少し下がった。

どこまで分かった?

結果は、このデータ、DCGAN、Swin Transformer、追加枚数と評価方法の組合せに基づく。ほかの生成法や分類器でも画像追加が無効だと結論づけるものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

医用画像データを増やすために敵対的生成ネットワーク(GAN)が使われることが増えているが、合成画像の追加が後続の分類に必ず役立つとは限らない。本研究では、分類器と評価用データを同じ条件に保ち、クラス別に学習した深層畳み込みGAN(DCGAN)による画像追加が脳腫瘍分類を改善するかを調べる。実験には、神経膠腫、髄膜腫、下垂体腫瘍、腫瘍なしの4クラス、計7200枚の脳MRI画像を用いた。各クラスで実画像1400枚を学習に、400枚を試験に取り分けた。実画像だけで学習したSwin Transformer分類器と、同じ実画像にDCGAN生成画像を各クラス500枚加えて学習した分類器を比較し、どちらも同じ取り分け済み試験データで評価した。 両モデルの全体正解率はともに96%で、マクロF1も実質的に変わらなかった。ROC-AUCは画像追加後に0.987から0.982へわずかに低下した。クラス別に見ると、誤分類の分布には小さな変化があったが、一貫した性能向上はなかった。実画像と合成画像の間のFIDは209.15~314.27で、この評価条件では分布に大きな差があることを示した。結果は、合成画像によるデータ拡張が医用画像分類を改善すると決めつけず、画像分布の忠実さと後続タスクでの有用性をともに評価すべきだと示唆する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generative adversarial networks (GANs) are increasingly used to augment medical imaging datasets, but synthetic images do not necessarily provide downstream classification benefits. This study investigates whether class-specific Deep Convolutional Generative Adversarial Network (DCGAN) augmentation improves brain tumor classification when the classifier and evaluation set are held constant. Experiments were conducted on 7,200 brain magnetic resonance imaging (MRI) scans across four classes: glioma, meningioma, pituitary tumor, and no tumor. For each class, 1,400 real images were used for training and 400 were reserved for testing. A baseline Swin Transformer classifier was trained using only the real training images and compared with a second model trained using the same real images augmented with 500 DCGAN-generated images per class. Both conditions were evaluated on the identical held-out test set. The two models achieved the same overall accuracy of 96%, while macro F1 remained effectively unchanged and ROC-AUC decreased slightly from 0.987 to 0.982 after augmentation. Class-level analysis showed small redistributions in errors rather than a consistent performance gain. FID values between real and synthetic images ranged from 209.15 to 314.27, indicating substantial distributional differences under the adopted evaluation setup. These results suggest that synthetic augmentation should not be assumed to improve medical image classification and should instead be evaluated for both distributional fidelity and downstream task utility.

arXiv ID: 2609.28508 / 要約の誤りについて