正解画像による調整を要しないベイズ型拡散画像復元
FB-GDM: Fully-Bayesian Guided Diffusion Models for High-Dimensional Linear Inverse Problems via Unsupervised Variational Inference
この論文をやさしく読む
ひとことで言うと
画像復元の際に、正解画像や雑音レベルを使った手動調整を減らす拡散モデルの手法です。
何に役立つ?
撮影条件や画像の種類が変わる逆問題で、観測データに忠実な復元を行う方法の検討に役立ちます。
この研究の面白いところ
画像復元を誘導する二つの重みを、各逆拡散ステップで観測値から推定します。
どこまで分かった?
要旨の定量結果はCelebA-HQを用いた逆問題実験です。すべての画像領域への一般化は示していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
拡散モデルは線形逆問題で強力な事前分布を与えるが、代表的な誘導手法であるDiffusion Posterior Sampling(DPS)と擬似逆行列誘導拡散モデル(ΠGDM)は、通常、課題ごとに正解データを使って調整するスカラーのハイパーパラメーターに依存する。本研究では、この調整をなくす完全ベイズ型の誘導拡散手法FB-GDMを提案する。ΠGDMのガウス近似から出発し、二つの精度パラメーター、すなわち逆分散に依存する条件付きスコアを閉形式で導く。一方はノイズ除去の近似、もう一方は観測の尤度に関係し、各逆拡散ステップで変分推論によって推定する潜在変数として扱う。分離可能な因子分解によって各更新の計算量は画素数に線形となり、画像の全解像度でもΠGDMを1回実行するのと同程度の費用で推論できる。必要な入力は観測値と順方向演算子だけで、雑音レベルも正解画像も不要である。CelebA-HQの逆問題実験では二つの結果を得た。第一に、観測値だけから推定した精度パラメーターによって、真の雑音レベルを与えられた通常設定のΠGDMより、演算子に応じて最大14 dB高い性能を示し、正解画像で調整したΠGDMの理想的な設定との差は0.1 dB以内だった。第二に、順方向演算子、雑音レベル、画像分布が変化しても、問題ごとに調整したΠGDMの性能に近く、DPSで観察された幻覚的な画像生成も生じなかった。一方、DPSは固定スケールで大きく性能が落ち、ΠGDMは新しい問題ごとに正解画像で再調整した場合にのみ競争力を保った。事前分布の学習に使われていない画像でも、データと事前分布の釣り合いを取り直すことで、固定された顔画像の事前分布による誘導では作り物の内容が生じ得る場面で、観測データへの忠実さを保った。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Diffusion models are powerful priors for linear inverse problems, but the reference guidance methods, Diffusion Posterior Sampling (DPS) and Pseudoinverse-Guided Diffusion Models ($\Pi$GDM), rely on scalar hyperparameters tuned per task, usually against the ground truth. We introduce FB-GDM, a fully-Bayesian guided diffusion method that removes this calibration step. Starting from the Gaussian approximation of $\Pi$GDM, we derive a closed-form conditional score that depends on two precision parameters (inverse variances), one associated with the denoising approximation and one with the observation likelihood, and treat them as latent variables inferred by variational inference at each reverse step. A separable factorization makes each update scale linearly with the number of pixels, so the inference stays tractable at full image resolution, at a cost comparable to one $\Pi$GDM run. FB-GDM requires neither the noise level nor the ground truth: its only inputs are the observation and the forward operator. Experiments on CelebA-HQ inverse problems establish two results. (i) The precision parameters, inferred from the observation alone, allow FB-GDM to outperform $\Pi$GDM at its nominal setting, even when the latter is given the true noise level, by up to 14 dB depending on the operator, and to match the ground-truth-calibrated $\Pi$GDM oracle within 0.1 dB. (ii) FB-GDM is robust when the forward operator, the noise level, or the image distribution changes: it stays close to a per-problem $\Pi$GDM oracle throughout and does not exhibit the hallucinations observed with DPS, whereas DPS substantially degrades at a fixed scale and $\Pi$GDM stays competitive only if it is re-tuned against the ground truth for each new problem. When the prior is applied to images outside its training set, this re-balancing between data and prior keeps FB-GDM faithful where a fixed face-prior guidance can otherwise hallucinate.
arXiv ID: 2609.29216 / 要約の誤りについて