arXiv論文メモ
新着一覧
cs.LG / cs.CV · 査読状況未確認

生成条件を考慮して拡散モデルの学習を安定させる

CARE: Condition-Aware Representation Regularization for Diffusion Models

Fengjia Guo, Zhuoyi Yang, Jie Tang

この論文をやさしく読む

ひとことで言うと

画像生成モデルが受け取るラベルや文章の似かよりを使い、学習中の特徴表現を整える手法。

何に役立つ?

考えられる用途は、クラス指定や文章指定の画像生成モデルで、画像品質と学習効率を改善すること。

この研究の面白いところ

追加の外部教師信号なしで既存の生成条件を使い、二種類の画像生成課題でFIDの改善を報告した。

どこまで分かった?

数値結果は要旨に示されたImageNetと文章からの画像生成の実験条件に関するもの。他の課題での効果はこの要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

拡散モデルの最近の進歩は、画像品質と学習効率を高める上で表現の正則化が重要であることを示している。しかし一般的な正則化手法は、生成対象を直接決めるラベルや文章といったモデル内の条件を見落としがちである。本研究は条件信号が特徴分布に及ぼす影響を示し、条件を考慮する表現正則化CAREを提案する。CAREは、条件の類似度に応じて特徴分布を動的に調節する、軽量で既存モデルに組み込みやすい枠組みである。明示的な位置合わせ損失や外部の教師信号を使わず、モデル内の条件信号によって表現空間を導き、似た条件の特徴をより密な集まりにする。 実験では、クラスから画像を生成する課題と文章から画像を生成する課題の両方で、視覚的な忠実度と収束の安定性が一貫して向上した。ImageNetでは40万学習ステップでFIDを19.08%下げ、3.5倍の高速化につながった。文章からの画像生成では20万反復でFIDを16.61%下げ、生成画像と文章指示の意味的な一致も改善した。既存の正則化手法とも組み合わせられ、さらに性能が向上した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent advances in diffusion models highlight the importance of representation regularization for improving sample quality and training efficiency. However, commonly used regularization methods often overlook the built-in conditions (such as labels or texts) which directly determine the generation target. In this work, we demonstrate how conditioning signals affect the feature distribution and introduce the CARE (Condition-Aware REpresentation regularization). CARE is a lightweight plug-and-play regularization framework that dynamically modulates feature distribution based on condition similarity. CARE leverages built-in conditioning signals to judiciously guide the representation space, promoting tighter feature clusters for similar conditions without relying on explicit alignment losses or external supervision. Empirically, CARE consistently improves both visual fidelity and convergence stability across both class-to-image and text-to-image tasks. On ImageNet, CARE achieves a 19.08\% reduction in FID in 400k training steps, leading to a 3.5$\times$ speed-up. When applied to text-to-image generation, CARE lowers FID by 16.61\% in 200k iterations and improves semantic alignment between generated samples and text prompts. Moreover, CARE can be seamlessly integrated with existing regularization methods, yielding additional performance gains.

arXiv ID: 2609.28561 / 要約の誤りについて