画像の再構成評価と生成品質が一致しない理由を調べる
Bridging Reconstruction and Generation: A Latent Distribution Perspective on Evaluation and Improvement
この論文をやさしく読む
ひとことで言うと
入力画像を上手に復元できるモデルが、必ずしも上手に新しい画像を作れるわけではない理由を、モデル内部の表現の違いから調べています。
何に役立つ?
画像生成モデルの部品を再構成品質だけで選んでよいかを診断し、生成時に近い内部表現を使ってデコーダーを調整する際に役立ちます。
この研究の面白いところ
再構成と生成を二つに分けるのではなく、その間を連続的に移ります。途中では元画像との対応が残るため、生成時に近づけつつ教師付きの調整が可能になります。
どこまで分かった?
GAR-FIDとgFIDの相関、生成品質の改善は実証結果として報告されていますが、相関係数や改善幅の数値は要旨にはありません。任意のモデルでの保証ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
潜在生成モデルでは、再構成品質が生成性能と相関すると考えられることが多い。しかし、再構成FID(rFID)と生成FID(gFID)の相関は弱い場合や、負になる場合すらある。本研究では、この食い違いを潜在分布の不一致に起因するものと捉える。再構成ではエンコーダーが作る潜在表現でデコーダーを評価するのに対し、生成では生成モデルが作る潜在表現に同じデコーダーを適用するためである。 この変化を特徴付けるため、生成を考慮した再構成(GAR)を導入する。GARは、エンコーダーの潜在表現にノイズを加え、生成モデルを通してノイズを除去してから復号することで、通常の再構成から生成へ向かう連続的な軌道を構成する。この軌道に沿ったデコーダーの振る舞いを調べることにより、エンコーダー由来の潜在分布から生成時の潜在分布への移行を観測・診断できるようにする。得られる軌道ベースの診断指標GAR-FIDは、多様なトークナイザーと規模で、gFIDと実証的に強い相関を示す。 重要なのは、中間のGAR潜在表現が、元画像との対応を保ちながら生成時の性質をより反映するようになる点である。これにより、完全に生成された潜在表現では得られない、対を成す教師情報が維持される。この対応を利用して中間のGAR潜在表現上でデコーダーを適応させると、さまざまなモデル規模で生成品質が一貫して改善する。全体として、潜在分布の不一致は、潜在生成モデルを評価・改善するための有用な視点となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In latent generative models, reconstruction quality is often assumed to correlate with generative performance. However, reconstruction FID (rFID) can exhibit weak or even negative correlation with generation FID (gFID). We attribute this discrepancy to a latent distribution mismatch: reconstruction evaluates the decoder on encoder-induced latents, whereas generation uses the same decoder on latents produced by the generative model. To characterize this shift, we introduce generation-aware reconstruction (GAR), which constructs a continuous trajectory from standard reconstruction toward generation by perturbing encoder latents with noise and denoising them through the generative model before decoding. GAR probes the decoder behavior along this trajectory, making the transition from encoder to generation-time latent distributions observable and diagnosable. The resulting trajectory-based diagnostic, GAR-FID, exhibits strong empirical correlation with gFID across diverse tokenizers and scales. Importantly, intermediate GAR latents become more generation-aware while preserving correspondence with their source images, thereby retaining paired supervision that is absent for fully generated latents. This correspondence enables decoder adaptation on intermediate GAR latents, consistently improving generative quality across model scales. Overall, latent distribution mismatch provides a useful perspective for evaluating and improving latent generative models.
著者のコメント
27 pages, 23 figures,and 15 tables
arXiv ID: 2609.24088 / 要約の誤りについて