arXiv論文メモ
新着一覧
cs.CV / cs.AI / cs.CL · 査読状況未確認

視覚言語モデルの成績は回答形式によって大きく変わる

What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit

Alfredo F. Frontera Del Valle

この論文をやさしく読む

ひとことで言うと

画像について同じことを尋ねても、位置や色をどう表記して答えさせるかで、視覚言語モデルの点数や順位が大きく変わる研究です。

何に役立つ?

視覚言語モデルのベンチマークを設計・比較する際、回答の表記方法と採点方法が能力評価をゆがめていないか点検する材料になる。

この研究の面白いところ

九つの位置を選ぶ課題で、同じ写真と位置でも英語名では68.5%、画素座標では20.0%と大きな差が出た。座標に誤った名前を付ける実験で、モデルがどの表記を使うかも調べている。

どこまで分かった?

表記への追従を見る方法では、試していない二つの形式の正答率を予測できなかった。結果にはモデルごとの差もあり、GPT-4oでは報告された写真課題の不利が見られなかった。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

視覚言語モデルのベンチマークは、選択肢を文字、色名、画素座標など、何らかの形式で示す。通常、その形式は中立だとみなされる。しかし本研究は、そうではなく、測定された限界がモデル自体ではなく回答の読み出し方法に由来する場合があることを示す。COCOの写真200枚について、Qwen3-VL-4Bが指定された物体の位置を九つから正しく選ぶ割合は、位置を英語名で示すと68.5%だが、同じ位置を画素座標で示すと20.0%だった。偶然に当たる確率は11.1%である。成績の低下は回答の選択肢が座標のときに生じ、質問文中に座標を与える場合の低下は3.5ポイントで、統計的に有意ではなかった。この差は4×4の格子でも、4ビットではなく8ビットの量子化でも、物体の大きさ、境界からの距離、カテゴリーごとのすべての区分でも保たれた。回答形式は、どのモデルが勝つかも左右する。英語名では同点の二つのモデルが、ある座標系では39ポイント、別の座標系では54ポイントの差を示し、優劣が逆転した。 色の課題では、それに対応できる四つの公開モデルのうち三つに形式による不利が見られた。写真の課題では、三つの公開モデルのうち二つに加え、Geminiでも解釈可能な回答に限ると11.1ポイントの不利が見られた(p=1×10⁻⁴)。GPT-4oにはその不利は見られなかった。モデルが座標を読めているかを調べるため、各座標に誤った名前を付け、モデルがどちらに従うかを記録した。色の選択肢を色相角で書くと、従う割合は偶然より低かった。一方、正規化した画素座標の形式には、偶然の4倍の割合で従った。これにより、モデルが使える形式と使えない形式を区別できたが、未試行の二つの形式における正答率までは予測できなかった。また、五つのモデルは同じ色相環に五通りの異なる名前を付けたため、固定された回答語彙もモデル間で中立ではない。この研究中、著者らの採点方法とモデルで回答の形式に対する認識が食い違ったため、能力のあるモデルを能力がないと測定した事例が五回あり、それぞれを報告している。これは論文が扱う現象を小さな形で示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Benchmarks for vision-language models offer their answer choices in some convention: a letter, a color name, a pixel coordinate. That convention is treated as neutral. We find it is not, and that the limits a benchmark reports can belong to the readout rather than to the model. On 200 COCO photographs, Qwen3-VL-4B picks the correct one of nine locations for a named object 68.5% of the time when the locations are given in English and 20.0% when the same locations are given as pixel coordinates. Chance is 11.1%. The cost arises when the answer options are coordinates; giving the model a coordinate in the question instead costs 3.5 points and is not significant. The gap holds on a 4x4 grid, under 8-bit rather than 4-bit quantization, and in every slice by object size, boundary distance and category. It also decides which model wins. Two models that tie under English names differ by 39 points in one coordinate system and by 54 in the other, in opposite directions. On the color task, three of the four open models capable of the task show the penalty; on photographs, two of three open models do, and so does Gemini, at 11.1 points on parseable answers (p = 1e-4). GPT-4o does not. To ask whether a model reads a coordinate at all, we attach the wrong name to each one and record which the model follows. Color options written as hue angles are followed below chance; a normalized pixel convention is followed at four times chance. This tells apart conventions a model can use from ones it cannot, though it did not predict accuracy on two untried conventions. Five models also name the same color wheel five different ways, so a fixed answer vocabulary is not neutral across models either. Five times during this work we measured a capable model as incapable because our scorer and the model disagreed about what an answer looks like. We report each case. They are the phenomenon in miniature.

著者のコメント

14 pages, 1 figure, 8 tables

arXiv ID: 2609.27408 / 要約の誤りについて