arXiv論文メモ
新着一覧
cs.CV · 掲載先の記載あり

合成医療画像で希少特徴の記憶と消失を監査

A Data-Interventional Framework for Auditing Privacy and Fairness in Generative Medical Imaging

Mischa Dombrowski and Bernhard Kainz

この論文をやさしく読む

ひとことで言うと

合成医療画像のAIが、珍しい特徴を別の画像にも生かすのか、元画像を丸ごと覚えるのか、特徴を消してしまうのかを調べます。

何に役立つ?

考えられる用途は、医療画像の合成データを共有する前に、記憶による漏えいリスクと、希少特徴の欠落を同時に点検する監査です。

この研究の面白いところ

人工的に注入した特徴を使い、何が保持・消失したかを制御された条件で調べます。珍しい条件付けが、画像の記憶を呼び出す鍵になるという関係も示しています。

どこまで分かった?

一般化しなかったという結果は、調べたモデルと条件付け、注入特徴に関する観測です。t′は記憶傾向の推定指標で、プライバシー保護の証明ではありません。軽減戦略の具体的な手順は要旨には示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

拡散モデルによる合成データ生成は、患者の機密記録を公開せずに医療画像データを共有する有望な方法である。しかし生成モデルには、プライバシーと公平性の間の根本的な緊張がある。まれな学習標本を記憶してプライバシーリスクを生む場合もあれば、少数の特徴を再現できず、不公平な合成分布を作る場合もある。従来研究は主として記憶と公平性を別々に扱っており、両者の相互作用は十分に理解されていない。本研究では、拡散モデルのプライバシーと公平性を系統的に解析する、データ介入型の枠組みを導入する。人手で注入したまれな画像特徴である合成解剖学的フィンガープリント(SAF)を、制御された検査用の手掛かりとして検討する。これにより、モデルが機微な属性を異なる個体へ一般化するのか、学習標本を記憶するのか、まれな信号を完全に抑えるのかを調べる。 複数の条件付けモダリティで一貫した挙動が観察された。モデルはフィンガープリントを忘れるか、それが含まれる画像全体を記憶するかのいずれかで、新しい画像へ一般化しなかった。標本の明示的な抽出が難しい大規模監査を支援するため、拡散過程の内部構造を利用して記憶しやすさを推定する指標t′も導入する。驚き度の異なる条件付け信号を比較し、条件の希少性と記憶挙動の明確な関係を明らかにする。驚き度の高い条件は記憶を増幅する検索キーとして働く一方、驚き度の低い条件は、その特徴が学習データに繰り返し現れていても、希少特徴を系統的に抑制する。本研究は、安全で公平な合成医療データ共有に向け、実行に移せる知見と具体的な軽減戦略を提供する。コードは https://github.com/MischaD/Privacy で公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
掲載先の記載あり

著者による掲載先の記載:Machine.Learning.for.Biomedical.Imaging. 2026 (2026)。出版社での独立確認は未実施です。

arXivで読むPDFDOI

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Diffusion-based synthetic data generation offers a promising route for sharing medical imaging data without releasing sensitive patient records. However, generative models face a fundamental tension between privacy and fairness: they may memorize rare training samples, leading to privacy risks, or fail to reproduce underrepresented features, resulting in unfair synthetic distributions. While prior work has largely focused on either memorization or fairness in isolation, their interaction remains insufficiently understood. In this work, we introduce a data-interventional framework to systematically analyze privacy and fairness in diffusion models. We discuss synthetic anatomical fingerprints (SAFs), rare and manually injected image features, as controlled probes to study whether models generalize sensitive attributes across identities, memorize training samples, or suppress rare signals entirely. Across multiple conditioning modalities, we observe a consistent behavior: models either forget these fingerprints or memorize the entire image in which they appear, but do not generalize them to novel images. To support large-scale auditing where explicit sample extraction is infeasible, we further introduce the indicator metric t', which estimates a model's susceptibility to memorization by exploiting the internal structure of the diffusion process. By comparing conditioning signals of varying surprisal, we reveal a clear relationship between conditioning rarity and memorization behavior. Highly surprising conditioning signals act as retrieval keys that amplify memorization, whereas low-surprisal conditioning signals systematically suppress rare features, even when these appear repeatedly in the training data. Our findings provide actionable insights and concrete mitigation strategies for safe and fair synthetic medical data sharing. Code is available at https://github.com/MischaD/Privacy.

著者のコメント

Accepted for publication at the Journal of Machine Learning for Biomedical Imaging (MELBA) https://melba-journal.org/2026:026

arXiv ID: 2609.26623 / 要約の誤りについて