arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

拡散モデルの誘導の強さを学習して生成手順を減らす

Learned End-to-End Guidance Schedules for Diffusion Models

Aneesh Barthakur, Mathias Niepert, Luiz F.O. Chamon

この論文をやさしく読む

ひとことで言うと

生成途中で条件に近づける力の強さを、時刻ごとに学習します。画質と条件の満足を両立しながら、拡散モデルの反復回数を減らす方法です。

何に役立つ?

画像補完や条件付き生成、偏微分方程式の問題で、限られた関数評価回数の中で性能を高めるために役立ちます。要旨では複数種類の課題で比較を行っています。

この研究の面白いところ

少数の例から誘導スケジュールを学び、さらにその学習自体に必要な逆伝播も近似で軽量化しています。生成時の費用とスケジュール学習の費用の両方を扱っています。

どこまで分かった?

同予算での優位性と10%の手順での同等性能は、評価した課題での報告です。すべての課題で同じ削減率を保証するという記載はなく、要旨には課題別の数値はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

拡散モデルは、マルチメディアや科学的な応用で広く使われる強力な生成の枠組みである。誘導付き拡散の方法では、微分可能な損失、すなわち誘導関数の勾配を、推論中のドリフト項として加えることで、生成に要件を課す。このドリフトの重みである誘導スケールは、データの品質と要件の充足とのトレードオフを左右する。両方の目標を達成するため、誘導付き拡散では小さな誘導スケールと長いサンプリング過程を使わざるを得ず、計算費用が高くなる。 本研究は、より少ないサンプリング手順でこれらの目標を達成するため、端から端まで学習する誘導スケジュール(LEEGS)を提案する。LEEGSは、少数の例に対する誘導関数を確率的勾配降下法で最小化して、時間に依存するスケジュールを学習する。誘導付きサンプリング全体を通した逆伝播は計算費用が高いため、LEEGSでは学習時間を4分の1に短縮する勾配の近似を用いる。画像の欠損補完、ノイズを含む画像の逆問題、顔の識別情報で誘導する生成、偏微分方程式の順問題と逆問題を含む多様な誘導課題で評価する。同じ計算予算、すなわち関数評価回数(NFE)が50回または100回の条件でベースラインを上回り、あるいは一定の誘導を使う方法のわずか10%の手順で同等の性能を達成する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Diffusion models are a powerful generative paradigm used across multimedia and scientific applications. Guided diffusion methods impose requirements on the generation by adding the gradient of a differentiable loss (the guidance function) as a drift term during inference. The weight of this drift (the guidance scale) is critical for the trade-off between data quality and requirement satisfaction. To achieve both of these goals, guided diffusion must resort to small guidance scales and lengthy sampling, incurring high computational costs. This work proposes learned end-to-end guidance schedules (LEEGS) to achieve these objectives with fewer sampling steps. LEEGS trains a time-dependent schedule by minimizing the guidance function over a small set of examples using stochastic gradient descent. Backpropagating through guided sampling is computationally expensive, so LEEGS uses an approximation of the gradient that cuts training time by a factor of 4. We evaluate LEEGS on diverse guidance tasks, including (a) image inpainting, (b) noisy image inverse problems, (c) face-ID-guided generation, and (d) forward and inverse PDE problems, outperforming baselines at equal budget (50 or 100 NFEs), or matching constant guidance with only 10% of the steps.

arXiv ID: 2610.01502 / 要約の誤りについて