arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

ほぼ同じ二枚の画像の小さな違いを説明し位置を示す

MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs

Chaoqian Mu, Wenhao Wu, Zichen Liang, Jiaxu Li, Lijun Wang, Yifan Wang, Huchuan Lu

この論文をやさしく読む

ひとことで言うと

二枚の画像にある一つの小さな違いを、言葉と画像上の位置の両方で答えられるか測った。

何に役立つ?

考えられる用途は、画像比較モデルの評価や、細かな変更の説明・位置特定を同時に改善すること。

この研究の面白いところ

違いの意味を捉え、空間的な手がかりを挟んでから説明する構成で、推論用トークンも約26%減らした点。

どこまで分かった?

要旨は改善の具体的な精度値を示していない。結果はMinCUなどの評価設定に基づく。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ほとんど同じ二枚の画像の細かな違いを見つけて説明する能力は、マルチモーダル大規模言語モデルにとって重要だが、十分に調べられていない。従来のベンチマークは、意味の比較か単一画像での位置特定を別々に評価することが多く、忠実な説明と物理的な位置特定の両方を求めない。本論文は、一組の画像が、物体の種類、属性、数、または空間的な位置に関する一つだけの最小変化で異なるベンチマークMinCUを提案する。モデルには、変化の説明、変化した領域の位置特定、変化した対象の特定を求める。 さらに、意味的な手がかりを用いる暗黙の空間アンカー法SG-ISAを提案する。この構造化された自己回帰法では、予測を「考える・位置を示す・説明する」の順に分ける。まず変化した概念の意味的な手がかりを予測し、離散的な空間アンカーを暗黙の位置特定の足場に用い、最後に変化の説明と対象領域の矩形を生成する。実験では、強力な非公開モデルや新しい推論モデルも、正確な説明と位置の矩形を同時に出すことに苦戦した。従来の思考過程を使う方法と比べ、SG-ISAで追加学習すると、位置特定精度と説明品質の双方が大きく改善し、推論用トークンの負担は約26%減った。この結果は、モデル規模だけに頼るより、途中段階に暗黙の空間的な仕組みを置くことが有効な可能性を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Localizing and describing fine-grained differences between near-identical images is a critical yet underexplored capability for multimodal large language models (MLLMs). Existing benchmarks largely assess semantic comparison or single-image grounding in isolation, without jointly requiring faithful description and physical localization. To bridge this gap, we introduce MinCU, a benchmark for grounded minimal-change understanding, where each sample consists of an image pair differing by a single atomic variation in object category, attribute, count, or spatial position, and models are evaluated on their ability to describe the change, localize the changed regions, and identify the changed entity. We further propose Semantic-Guided Implicit Spatial Anchors (SG-ISA), a structured autoregressive method that decomposes prediction into a Think-Locate-Describe sequence. SG-ISA first predicts a semantic cue for the changed concept, then uses discrete spatial anchors as an implicit localization scaffold, and finally generates the change description together with the grounding box. Experiments reveal that even the strongest closed-source MLLMs and recent R1-style reasoning models struggle on MinCU, with most failing to jointly produce accurate descriptions and grounding boxes. Compared to the previous chain-of-thought method, fine-tuning with SG-ISA yields substantial joint improvements in grounding accuracy and description quality while reducing reasoning-token overhead by approximately 26%. These results suggest that an implicit intermediate spatial interface can be more effective than relying solely on model scale for grounded dual-image understanding.

著者のコメント

25 pages, 10 figures, 15 tables

arXiv ID: 2609.23336 / 要約の誤りについて