arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

画像位置合わせのずれが合成CTの学習評価に与える影響

When Misalignment Becomes Supervision: Structured Label Noise in Supervised Synthetic CT Generation

Valentin Boussot, Cedric Hemon, Caroline Lafond, Jean-Claude Nunes, Jean-Louis Dillenseger

この論文をやさしく読む

ひとことで言うと

MRIやCBCTからCTを作るモデルでは、画像位置合わせの残るずれが学習と評価の両方をゆがめることを調べています。

何に役立つ?

合成CTの開発や評価で、画素単位のスコアに加え解剖学的な構造を確認する理由を示します。臨床上の成果を直接検証した報告ではありません。

この研究の面白いところ

学習と評価の位置合わせ方式が一致するだけでスコアが上がり、CTだけの対照実験でも大きな指標誤差が出ました。

どこまで分かった?

評価は五つの解剖学的領域の1,784人の対応画像に基づきます。知覚損失による改善は画像と下流領域分割で示され、患者転帰は要旨に記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

教師ありの合成CT画像生成では、位置合わせした参照CT画像を正解とする画素単位の回帰として、学習と評価を行うことが多い。しかし実際には、MRIとCT、またはCBCTとCTの画像対を位置合わせしてもずれが残る。この残差は独立な画素値の雑音ではなく、空間的につながった幾何学的な不一致であり、構造を持つラベル雑音として働く。 本研究は、五つの解剖学的領域にわたる1,784人の対応画像を用い、この位置合わせに由来する偏りが、MRIからCTおよびCBCTからCTへの教師あり画像生成に与える影響を調べた。画素単位のスコアは、学習用の正解画像を作る位置合わせと、評価に用いる位置合わせの整合性に強く依存した。両者の方式が一致するとモデルのスコアが最も良くなり、ネットワークが位置合わせ処理の幾何学的な慣例を部分的に学び、標準的な指標がそれを高く評価していることが示された。解剖学的な整合性が高い位置合わせで学習すると予測のばらつきが減り、分布外のデータに対する頑健性が向上した。またCTだけを使う対照実験では、位置合わせだけで、主要なコンペティション提出結果と同程度の指標誤差が生じた。 画素単位の教師信号の限界を緩和するため、事前学習したSegment Anythingのエンコーダの特徴空間で計算する知覚損失を導入した。平均絶対誤差だけの目的関数やVGGを使う目的関数と比べて、下流の領域分割が改善し、より鮮明で構造的に整った合成CTが得られた。位置合わせが不完全なときには知覚指標と画素単位の指標は一致せず、評価用画像の幾何学的な配置が信頼できるときには一致した。これらの結果は、位置合わせに由来する偏りが教師あり合成CT生成の重要な交絡要因であり、画素値の一致だけでなく解剖学的構造を重視した評価を組み合わせる必要があることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Supervised synthetic CT (sCT) generation is commonly trained and evaluated as voxel-wise regression against registered reference CT images. In practice, MRI-CT and CBCT-CT pairs are aligned through registration procedures that leave residual misalignments. These residuals are not independent intensity noise but spatially coherent geometric discrepancies that act as structured label noise. We investigate how this registration-induced bias affects supervised MRI-to-CT and CBCT-to-CT synthesis on 1,784 paired patients covering five anatomical regions. Voxel-wise scores strongly depend on the consistency between the registration used to build the training targets and the one used for evaluation: models score best when both conventions match, showing that networks partly learn the geometric convention of the registration pipeline and that standard metrics reward it. Training on more anatomically consistent registrations reduces prediction variability and improves out-of-distribution robustness, and CT-only controls show that registration alone produces metric errors in the range of top challenge submissions. To mitigate the limits of voxel-wise supervision, we introduce a perceptual loss computed in the feature space of a pretrained Segment Anything encoder. Compared with MAE-only and VGG-based objectives, it improves downstream segmentation and yields sharper, more structurally coherent sCT. Perceptual and voxel-wise metrics disagree under imperfect alignment and agree when the evaluation geometry is reliable. These results identify registration-induced bias as a central confounder in supervised sCT generation and argue for complementing voxel-wise agreement with anatomy-oriented evaluation criteria.

著者のコメント

23 pages, 4 figures, 19 tables

arXiv ID: 2609.29387 / 要約の誤りについて