arXiv論文メモ
新着一覧
cs.CV / cs.CL / cs.LG · 査読状況未確認

画像の美しさを複数のAIで評価する際の採点項目の効果

Feed the Panel Dimensions, Not Verdicts: Rubric-Decomposed Fusion of Vision-Language Aesthetic Judges

Amit Jadhav, Shaurya Beriwala, Beomjin Kim

この論文をやさしく読む

ひとことで言うと

画像の美しさを複数の視覚言語モデルで判定する際、総合点だけでなく評価項目ごとの点数を統合する効果を調べた。

何に役立つ?

画像の審美評価システムで複数モデルの情報を組み合わせる設計に役立つ可能性がある。EVA では改善したが、PARA では同等か小幅な低下だった。

この研究の面白いところ

単純な総合判定の寄せ集めでは最良の単一モデルを有意に上回らず、五つの項目別点数を使うと EVA の10組すべてで上回った。

どこまで分かった?

改善には数百件のラベルと EVA で4.8倍の API 呼び出しが必要で、ラベルはデータセット間で転用できなかった。事前登録で失敗した設定もある。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚言語モデルは画像の美しさを追加学習なしで判定するために使われ、複数モデルによる判定が信頼性を高める方法として勧められてきたが、その証拠は十分ではない。人間が採点した EVA と PARA の二つのデータセットで調べると、画像を総合評価するモデルを複数集めても、判定を平均する場合も学習した結合器で統合する場合も、最良の単一モデルを有意には上回らなかった。そこで、人間が書いた固定の評価基準の五項目について各モデルに画像を採点させ、その項目別の点数と総合判定を、モデル群をまたいで交差検証の訓練外予測を用いた結合器で統合した。項目別の点数は名称に対応した情報を測っており、人間の総合点の効果を除いても、30のモデルと属性の組み合わせのうち28で、項目を尋ねるプロンプトは総合判定用プロンプトより属性固有の情報を多く含んだ。統合結果は EVA の三つのモデル群からなる10通りの組み合わせすべてで最良の単一モデルを上回った。最も強い組み合わせでは Spearman の相関係数が0.07、事前指定の組み合わせでは0.10高く、20通りの分割で平均するとそれぞれ0.06と0.07高かった。パネル平均との比較では、主要な検定で EVA の設計用データ上の差は0.118だった。PARA では Spearman の相関係数は同等、Kendall の tau-b では小さく有意でない低下となり、ここでは一つのモデルが人間による評点の雑音による上限の85%をすでに捉えていた。改善は特徴量の列を増やしただけでは説明できず、同じ繰り返しから得た総合判定だけの列を同数与えても再現できなかった。改善には数百件のラベルが必要で、それらはデータセット間で転用できず、EVA では API 呼び出しが4.8倍になった。著者らは対応のあるブートストラップと Kendall の tau-b による結果に加え、事前登録での失敗と性能が下がった設定も報告する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Vision-language models (VLMs) are deployed as zero-shot judges of image aesthetics, and panels of several models are recommended, on thin evidence, as the way to make such judges reliable. On two human-rated datasets, EVA and PARA, we find that a panel of holistic judges never significantly beats its best member, whether the verdicts are averaged or fused by a learned combiner. What a panel is worth depends on what it is fed. We therefore have each model score each image on the five dimensions of a frozen, human-written rubric and fuse those scores, alongside each model's verdict, across model families with an out-of-fold combiner. The dimension scores measure what their labels claim: with the overall human score partialled out, a dimension prompt carries more attribute-specific information than the holistic prompt in 28 of 30 model-attribute cells. Fused, they beat the best single VLM in all ten three-family panels on EVA (against that best single model, +0.07 Spearman rho for the strongest trio and +0.10 for the pre-declared one, and +0.06 and +0.07 when averaged over twenty fold partitions; against the panel mean, the primary test gives +0.118 on its EVA design set), and on PARA they reach parity under Spearman rho and a small, non-significant loss under Kendall tau-b, where one model already captures 85% of the human noise ceiling. It is not a feature-count artefact: giving the same combiner an equal number of pure holistic columns, split from the same repetitions, does not reproduce it. The gain costs a few hundred labels, which do not transfer between datasets, and 4.8x the API calls on EVA; we report it with paired bootstraps and Kendall tau-b, alongside a failed pre-registration and the configurations that lost.

著者のコメント

19 pages, 7 figures

arXiv ID: 2609.27110 / 要約の誤りについて