学習時と分布が違うデータに対する拡散モデルの頑健性
Robustness of Diffusion Models under Distribution Shift
この論文をやさしく読む
ひとことで言うと
学習時と実際のデータ分布が違うとき、拡散モデルの推定誤差がどこから生じるかを理論的に分けた研究です。
何に役立つ?
分布変化を受ける拡散モデルの保証を考える際、データ不足と分布のずれの寄与を別々に評価する基準になります。
この研究の面白いところ
分布のずれによる費用がワッサースタイン半径の二乗に比例し、その依存性が最適であることを示し、半径を知らない推定器も構成しています。
どこまで分かった?
主な結果は要旨に記されたオルンシュタイン・ウーレンベック拡散と分布の仮定での理論保証です。実データの実験結果は要旨にありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
スコアに基づく拡散モデルは、実際のデータ分布が学習時の分布と異なり得る状況でも検討されるようになっているが、既存の理論的保証は主に分布の変化がない場合に集中している。本研究は、基準分布のワッサースタイン摂動の下での頑健なスコア推定を調べる。オルンシュタイン・ウーレンベック拡散では、頑健な推定が、基準分布を学習する統計的な費用と、分布変化そのものに由来する費用という二つの基本要素に分かれることを示す。後者はワッサースタイン半径の二乗に比例し、この依存性はミニマックスの意味で最適である。分布変化の半径を知らなくても、この頑健なミニマックス速度を達成する、明示的な有限標本推定器を構成する。 基準分布が未知の低次元部分空間上にある場合、統計的な項はその内在次元に適応するが、分布変化の費用は変わらない。最後に、正の時刻から始める逆方向のサンプリングも同じ分解に従うことを示し、KLダイバージェンスに関して対応するミニマックス保証を得る。これらの結果は、有限のデータ、内在次元、分布変化がスコア型拡散モデルの頑健性にどう影響するかを特徴付ける。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Score-based diffusion models are increasingly considered in settings where the underlying data distribution may differ from the training distribution, yet existing theoretical guarantees largely focus on the no-shift setting. In this work, we study robust score estimation under Wasserstein perturbations of a reference distribution. For the Ornstein--Uhlenbeck diffusion, we show that robust estimation decomposes into two fundamental components: the statistical cost of learning the reference distribution and the intrinsic cost of distribution shift. The latter scales quadratically with the Wasserstein radius, and this dependence is minimax optimal. We construct an explicit finite-sample estimator achieving the resulting robust minimax rate without knowing the shift radius. When the reference distribution lies on an unknown low-dimensional subspace, the statistical term adapts to the intrinsic dimension while the shift cost remains unchanged. Finally, we show that the same decomposition governs positive-time reverse sampling and obtain matching minimax guarantees in KL divergence. Together, these results characterize how finite data, intrinsic dimension, and distribution shift affect the robustness of score-based diffusion models.
著者のコメント
13 pages, 2 figures
arXiv ID: 2609.27546 / 要約の誤りについて