arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

拡散モデルで画像が分かれる時点から知覚距離を作る

FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

Jaihyun Lew, Mingi Jung, Minjun Park, Wooseok Song, Sungroh Yoon

この論文をやさしく読む

ひとことで言うと

拡散モデルで二枚の画像が生成過程のいつ分かれたかを使い、人手なしで知覚距離の教師ラベルを作る研究。

何に役立つ?

画像品質評価モデルの訓練データを、大量の人手評価に頼らず作る方法の参考になる。

この研究の面白いところ

大まかな構造が決まる前に分かれた画像ほど遠いと見なし、任意の画像対を比べられる点ごとのラベルを作る。

どこまで分かった?

要旨は複数ベンチマークでの優位を述べるが、各ベンチマークの具体的な差や全ての種類の画像への汎化は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

参照画像を使う画像品質評価指標は、二枚の画像間の知覚的な距離を人間の感じ方に近づけようとする。最近の指標は、人間の視覚系の働きを学ぶために人手で付けた評価データに大きく依存する。画像ごとに品質を数値で付ける平均意見得点(MOS)は注釈には便利だが、大規模に集めるには費用が高く、人間の判断の不一致による雑音もある。二択強制選択(2AFC)の画像対ラベルは信頼性と効率のため広く使われるが、対の相対比較しか捉えない。本研究は、人間の注釈を使わず、画像対の点ごとの知覚距離ラベルを生成する完全自動のデータ作成手順を提案する。拡散モデルでは初期の段階で画像の大まかな構造が作られ、後の段階で細部が作られるという生成の動きを、知覚距離の代用にする。生成過程の早い時点で枝分かれした画像は大まかな構造しか共有せず知覚的に遠いが、遅く枝分かれした画像は細部だけが違う。拡散の軌跡が人間の視覚系とよく合うことを示し、この分岐時点FoMoを参照画像に根ざした距離ラベルとして、画像品質評価指標の学習に用いる。点ごとのラベルは任意の画像対を横断して比較でき、情報量の多い学習目的を可能にする。異なる基盤構造での幅広い実験は、このデータ作成手順の有効性を確認し、複数のベンチマークで人手注釈データを上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance between a pair of images. To learn how the human visual system (HVS) operates, recent reference-based IQA metrics heavily rely on human-annotated data. Mean opinion score (MOS)-based pointwise scoring, which assigns a scalar quality value per image, is preferable for annotation but is prohibitively expensive to collect at scale and is known to be noisy due to inconsistent human judgments. As an alternative, two-alternative forced choice (2AFC) pairwise labels have gained popularity due to their reliability and efficiency, but they capture only relative comparisons between pairs. In this paper, we propose a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation. Our approach exploits the generative dynamics of diffusion models as a perceptual distance proxy, where the coarse structure of an image is generated in the early timesteps and the fine details are generated in the later timesteps. Images that fork early in the generation process share only coarse structure and are perceptually far apart; images that fork late differ only in fine detail. We demonstrate that the diffusion trajectory aligns well with the human visual system, and use this forking moment, FoMo, as a reference-grounded distance label to supervise the training of a reference-based IQA metric. The pointwise labels, which support universal comparison between arbitrary image pairs, enable an information-rich training objective. Extensive experiments across diverse backbone architectures confirm the effectiveness of our generation pipeline, outperforming human-annotated datasets in multiple benchmarks.

arXiv ID: 2609.25716 / 要約の誤りについて