arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

医用画像トークナイザーの再構成・生成・記憶を比較評価

What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization

Niklas Bubeck, Yundi Zhang, Vasiliki Sideri-Lampretsa, Julian McGinnis, Jiancheng Yang, Daniel Rueckert, Jiazhen Pan

この論文をやさしく読む

ひとことで言うと

医用画像を生成モデル用に圧縮する部品について、画像の復元だけでなく生成・分類・訓練データの記憶まで横断して比べます。

何に役立つ?

自然画像向けの評価をそのまま当てはめず、医用画像の用途に合わせた圧縮表現を選ぶ判断材料になります。

この研究の面白いところ

コードブックをほぼ使い切っていても潜在空間全体の利用は少ないことや、再構成と生成の性能が強く相関することを報告しています。

どこまで分かった?

30構成・12データセットなどの評価範囲内の結果です。訓練データの記憶が軽度だったことは、患者情報の漏えいが起こらない保証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

潜在拡散モデルは現在、医用画像生成の主流となっている。この種の処理系はすべて、画像を生成処理が扱う潜在コードに圧縮するトークナイザーを基盤としている。そのためトークナイザーの選択は、再構成の忠実度と生成品質から、後段の分析で利用できる表現まで、あらゆる後段タスクの上限を定める。しかし医用画像の処理系では、自然画像用のトークナイザーの振る舞いがそのまま通用するという仮説のもと、それらを日常的に利用している。この仮定は、データセットが桁違いに小さく、画像間のばらつきもはるかに小さい医用画像の領域では、これまで検証されていない。 そこで、10モデル系列の30構成を、12データセットと三つの圧縮率で評価する、医用画像トークナイザーの体系的な検証を示す。対象は再構成、生成、潜在空間の幾何、後段の分類、記憶にわたる。結果として、(1)自然画像についての先行報告と異なり、画像再構成と生成の性能は強く相関する、(2)現在のトークナイザーはコードブックのほぼ全項目を使うが、それでも潜在空間の大部分は未使用のままである、(3)訓練集合の記憶は軽度であり、潜在空間をより強く圧縮するとさらに抑えられる、(4)離散量子化は後段の分類性能をおおむね維持できるが、参照表を使わない方式が主な例外となる、ということが分かった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the latent codes for image generation to operate on. Thereby, the tokenizer choice bounds every downstream task from reconstruction fidelity and generation quality to the representations available for downstream analysis. Yet, medical imaging pipelines routinely utilize tokenizers from natural imaging on the hypothesis that their behavior carries over. However, this is an assumption never tested in the medical imaging regime, where datasets are orders of magnitude smaller and images exhibit far lower inter-sample variance. We present a systematic evaluation of medical image tokenizers evaluating thirty configurations across ten model families on twelve datasets at three compression factors, spanning reconstruction, generation, latent geometry, downstream classification, and memorization. We find that (1) performance on image reconstruction and generation strongly correlate, unlike prior reports on natural images; (2) modern tokenizers use nearly all of their codebook entries, but still leave most of the latent space unused; (3) training-set memorization is mild and is further suppressed by stronger latent space compression; and (4) discrete quantization can largely preserve downstream classification, with lookup-free schemes being the main exception.

arXiv ID: 2609.24691 / 要約の誤りについて