arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

同じ拡散モデルの難易度差を使って画像生成を学習する

SALD: Self-Referenced Advantage Learning for Diffusion Models

Aryan Das, Surjo Dey, Koushik Biswas, Swalpa Kumar Roy, Moloud Abdar, Arnab Bhattacharya, Vinay Kumar Verma

この論文をやさしく読む

ひとことで言うと

同じ画像をノイズが少ない場合と多い場合で処理し、その難しさの差を使って拡散モデルの学習を調整する方法です。

何に役立つ?

外部の教師モデルや推論手順の変更を加えずに生成品質を改善する学習方法として検討できます。複数の構成とデータセットで改善を報告しています。

この研究の面白いところ

容易な経路をそのまま正解としてまねるのでなく、難しい経路との誤差差を重みに変えます。難しさの履歴や周波数別の残差も同じモデル内で利用します。

どこまで分かった?

追加の学習可能なパラメータが不要であることと、学習計算量が増えないことは別です。要旨には計算費用や具体的な品質改善量、データセット名は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

近年の言語モデル適応の研究では、実演やフィードバックを加えた文脈で自らの振る舞いを評価し、生徒モデルが学習したパラメータに基づいて動く教師ネットワークの助けを借りることで、単一のモデルから有益な学習信号を得られることが示されている。この内部参照の原理に着想を得て、拡散モデルが外部の実演や教師ネットワークなしに、自己参照による学習信号を見いだす方法を検討する。 同じモデルを使い、各画像・キャプション対を2段階のノイズ水準で評価する自己参照型の学習枠組みSALDを導入する。容易な低ノイズの経路は勾配を追跡せずに評価して参照を与え、難しい高ノイズの経路からは学習用の勾配を得る。SALDは容易な経路の予測を直接蒸留するのではなく、2経路の誤差の差を使って難しい経路の目的関数を調整する。提案するAdvantage-Guided Diffusion(AGD)は、この相対誤差を微分可能な標本単位の重みに変換する。Temporal Advantage Memory(TAM)は学習を通じて相対的な難しさを蓄積し、その後の2つのノイズ水準の間隔を適応させる。さらにSpectral Advantage Decomposition(SAD)は、両経路の残差のパワースペクトルを比較し、周波数から導かれる微分可能な潜在要素ごとの重みを構成する。 すべての構成要素は単一のモデルパラメータ集合を共有する。学習時にも推論時にも外部の教師ネットワークや追加の学習可能なパラメータを必要とせず、推論手順の変更も不要である。複数の構成とデータセットでの実験により、生成品質の一貫した改善を示し、構成要素ごとの除去比較によって提案要素の寄与を定量化する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent work on language-model adaptation has shown that single models can obtain informative training signals by evaluating their behavior in demonstrationor feedback-augmented contexts, with the help of a teacher network, which is driven by the student's learned parameters. Inspired by this internal-reference principle, we investigate how diffusion models can identify self-referenced training signals without external demonstrations or teacher networks. We introduce SALD, a self-referenced training framework that evaluates each image-caption pair at two noise levels using the same model. The easier, lower-noise path is evaluated without gradient tracking to provide a reference, while the harder, higher-noise path provides the training gradient. Rather than directly distilling the easy-path prediction, SALD uses the difference between two path errors to adapt the hardpath objective. The proposed Advantage-Guided Diffusion (AGD) converts this relative error into a differentiable sample-level weight. Temporal Advantage Memory (TAM) accumulates relative difficulty across training and adapts the future gap between the two noise levels. Spectral Advantage Decomposition (SAD) further compares the residual power spectra of the two paths and constructs a differentiable, frequency-derived latent-element weight. All components share a single set of model parameters, requiring neither an external teacher network nor additional trainable parameters during training or inference, and no modification to the inference procedure. Experiments across multiple architectures and datasets demonstrate consistent improvements in generation quality, while component-wise ablations quantify the contributions of the proposed components.

arXiv ID: 2610.01496 / 要約の誤りについて