arXiv論文メモ
新着一覧
cs.LG / cs.CV · 査読状況未確認

拡散モデルの確率的生成を一つの速度場で表す

Mean Velocity Matching: Rethinking Generative Dynamics in Diffusion Models

Yunhong Zhang, Changjie Cao, Zhihua Zhang, Bingli Liu, Zongjie Cao, Zongyong Cui, Ying Yang

この論文をやさしく読む

ひとことで言うと

画像生成の拡散モデルで、確率的な逆向き生成を一つの学習済み速度場から行う方法を提案しました。

何に役立つ?

確率的・決定論的なサンプリング方式を同じ表現で比較し、必要な計算回数に応じて選ぶ検討に役立ちます。

この研究の面白いところ

追加のスコア推定を省きつつ、時刻0近くで学習目標が発散する問題を尺度調整で扱っています。

どこまで分かった?

ImageNet実験の具体的なFIDと関数評価回数は、提供された要旨では未展開のマクロになっており確認できません。性能比較は要旨の定性的な記述にとどめます。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

拡散モデルの確率的な生成過程について、予測量の表し方を調べる。速度に基づく既存の生成モデルは一つの輸送場を学習すればよいという単純さを持つが、標準的な形式は決定論的であり、確率的な拡張では通常、追加のスコア情報や速度からスコアを再構成する段階が必要になる。本研究は、一つの場の予測を保ったまま確率的な逆向きの過程を直接扱うMean Velocity Matching(MVM)を導入する。 MVMは、元のデータへ戻す向きの速度(x₀−xₜ)/tの条件付き期待値が、逆向きの確率微分方程式のドリフトを直接構成するようなガウス摂動過程を作る。これにより、スコアを別に推定・再構成しなくても、一つの学習済みの場で確率的な逆過程を表せる。この速度をそのまま回帰するとtが0に近づくと発散するため、MVMは逆向きの過程を保ちながら学習目標を有界にする、√tで尺度を調整した表現も導入する。同じ学習済みの場から決定論的な確率流常微分方程式も導かれ、確率的・決定論的なサンプリングを共通の定式化で調べられる。 Transformerを用いた生成モデルの実験では、ImageNetの32×32画像と256×256画像で生成品質を評価した。要旨中のFIDと関数評価回数の欄は具体的な数値ではなく未展開のマクロ表記になっているため、数値は確定できない。制御した確率微分方程式と常微分方程式の比較では、関数評価回数が非常に少ないときは常微分方程式が優れ、十分な評価回数があると確率的な逆過程のFIDが低かった。MVMは、確率的な逆向き生成を一つの場で直接表しながら、競争力のある生成品質を維持すると報告する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

This work studies prediction parameterization for stochastic generative dynamics in diffusion models. Existing velocity-based generative models provide the simplicity of learning a single transport field, but their standard formulation is deterministic, whereas stochastic extensions generally require additional score information or an intermediate velocity-to-score reconstruction. To retain single-field prediction while directly supporting stochastic reverse dynamics, this paper introduces Mean Velocity Matching (MVM). MVM constructs a Gaussian perturbation process for which the conditional expectation of a restoration-oriented velocity, $(x_0-x_t)/t$, directly forms the reverse-SDE drift. Consequently, a single learned field is sufficient to parameterize the stochastic reverse process without separately estimating or reconstructing the score. Because direct regression of this velocity becomes unbounded near $t=0$, MVM further introduces a $\sqrt{t}$-scaled parameterization that preserves the reverse dynamics while yielding a bounded training target. The same learned field also induces a deterministic probability-flow ODE, enabling stochastic and deterministic sampling to be studied within a unified formulation. Experiments with Transformer-based generative models achieve an FID of $\MVMImageNetThirtyTwoFID$ at \MVMImageNetThirtyTwoNFE\ NFE on ImageNet $32\times32$ and $\MVMImageNetTwoFiftySixFID$ at \MVMImageNetTwoFiftySixNFE\ NFE on ImageNet $256\times256$. Controlled SDE--ODE comparisons further show that the ODE performs better under very low NFE, whereas the stochastic reverse process achieves lower FID when sufficient function evaluations are available. These results demonstrate that MVM provides a direct single-field parameterization of stochastic reverse dynamics while maintaining competitive generation quality.

arXiv ID: 2609.25444 / 要約の誤りについて