新しい画像品質項目を評価するための複数エージェント手法
An Evolutionary Agentic Approach for Open-ended Image Quality Perception
この論文をやさしく読む
ひとことで言うと
画像の新しい品質項目を評価する際、具体的な視覚的問いを作って全体印象に引きずられた採点を抑える手法である。
何に役立つ?
物理的な妥当性や文字の正しさなど、既存の固定した評価項目では扱いにくい画像品質を評価する方法の検討に役立つ。
この研究の面白いところ
項目ごとに人が注釈した画像四枚で評価手順を較正し、全体印象による上書き率を44.4%から8.6%へ下げた。
どこまで分かった?
要旨では多様な評価条件での成績を述べるが、すべての新しい品質項目に一般化する保証や実運用での結果は記載されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
生成モデルの発展に伴い、画像品質評価は従来の忠実度だけでなく、物理的なもっともらしさや画像中の文字の正しさなど、新しい観点に広がっている。しかし既存の画像品質評価モデルは固定した定義と大量の教師データに依存し、自由に追加される知覚上の評価項目へ拡張しにくい。著者らは、未知の項目を採点するときに一般的な品質の先入観を再利用し、採点の誤りや順位の逆転を起こす、全体印象への偏りを重要な制約として特定する。これに対し、学習を必要としない複数エージェントの枠組みPACE(Perceptual Agentic Collaborative Evolution)を提案し、新しい画像品質評価を明示的な評価手順の作成として捉える。対象の評価項目を与えると、複数のエージェントが協調し、検証可能な視覚質問応答の問いからなる評価手順を構築する。これにより全体的な印象ではなく具体的な画像上の証拠に評価を基づかせる。できた手順は項目ごとに人が注釈した四枚の画像だけで較正し、二つの経路からなる採点の仕組みでモデルの知覚を人間の採点尺度に合わせる。従来型の画像品質評価、構造の忠実度、文脈を考慮した美的評価、新たに定義した評価項目で、PACEは基盤となるマルチモーダル大規模言語モデルを一貫して改善し、多様な評価条件で競争力のある性能を示した。全体印象による上書き率HORは44.4%から8.6%へ低下した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Generative models are rapidly expanding image quality assessment (IQA) beyond traditional fidelity factors to emerging dimensions such as physical plausibility and text-rendering correctness. However, existing IQA models rely on fixed definitions and heavy supervision, making them difficult to extend to open-ended perceptual dimensions. We identify holistic bias as an important limitation: when scoring an unseen dimension, models reuse generic quality priors, leading to scoring errors and rank inversion. To address this, we propose PACE (Perceptual Agentic Collaborative Evolution), a training-free multi-agent framework that formulates open-ended IQA as explicit protocol construction. Given a target dimension, PACE uses collaborative agents to construct an evaluation protocol composed of verifiable Visual Question Answering (VQA) probes, grounding evaluation in concrete visual evidence rather than holistic impressions. The resulting protocol is calibrated using only four human-annotated images per dimension, while a dual-track scoring mechanism aligns model perception with human scoring scales. Across traditional IQA, structural fidelity, context-aware aesthetics, and newly defined open-ended dimensions, PACE consistently improves its MLLM backbone, achieving competitive performance across diverse IQA settings, and reduces the Holistic Override Rate (HOR) from 44.4\% to 8.6\%.
arXiv ID: 2609.22942 / 要約の誤りについて