高解像度画像で探すべき証拠を質問文から明示する
Pay More Attention To Text In High-Resolution MLLMs
この論文をやさしく読む
ひとことで言うと
画像を詳しく見る前に、質問に答えるには何を探す必要があるかを文章で具体化する方法です。
何に役立つ?
考えられる用途は、高解像度画像の質問応答で必要な場所や細部を見つけやすくすることです。追加学習を必要としない方法として提案されています。
この研究の面白いところ
探索量や証拠の幾何条件をそろえた比較を行い、証拠の指定自体がもたらす効果を調べています。
どこまで分かった?
改善率はいずれも相対値で、正答率の増加幅をパーセントポイントで示したものではありません。最先端性能は著者らが評価したベンチマークでの報告です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
高解像度画像を扱うマルチモーダル大規模言語モデル(MLLM)の失敗は、視覚上の問題とされることが多い。そのため、細部の証拠を取り出したり干渉を抑えたりするために、拡大や切り出しなどの視覚的な介入が行われる。しかし近年の研究は、関連する視覚的証拠が中間表現にすでに符号化されていることを示唆しており、視覚側の改善だけでは不十分である。そこで、残る障害は視覚探索を導く文章にあるのではないかという疑問が生じる。 本研究では、回答を得るために書かれた質問が、場所を特定するのに必要な視覚的証拠を必ずしも指定しないという、これまで見落とされていた言語上の障害を特定する。この不一致に対処するため、最終的な推論には元の質問を保持しつつ、補完的な証拠の仕様を導く、学習不要の変換器EviSpecを導入する。さらに、証拠の指定と位置特定の役割を切り分ける、条件をそろえた対照実験で検証する。 探索予算を固定すると、構造化した証拠の仕様は一般的な要求より8.6%の相対改善をもたらす。証拠の幾何的条件をそろえると、EviSpecが位置を特定した証拠はランダムな証拠より14.8%の相対改善をもたらす。これらの対照実験は、単に視覚情報へのアクセスを増やすのではなく、探す証拠を指定することの効果を切り分ける。5つのMLLMすべてで、EviSpecは3つのベンチマークそれぞれの対応する基準手法を一貫して改善し、V*Bench、HR-Bench-4K、HR-Bench-8Kでの平均相対改善はそれぞれ10.4%、8.8%、12.4%だった。高解像度の推論に加え、視覚質問応答およびハルシネーションを重視したベンチマークでも最先端の性能を達成する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Failures of high-resolution MLLMs are commonly attributed to a visual problem, motivating zooming, cropping, and related visual interventions to recover fine-grained evidence or suppress interference. Yet recent studies suggest that relevant visual evidence is already encoded in intermediate representations, indicating that visual-side improvements alone insufficient. This raises a natural question: does the remaining bottleneck lie in the text that guides visual search? We identify a previously overlooked linguistic bottleneck: questions formulated for answering do not necessarily specify the visual evidence required for localization. To address this mismatch, we introduce EviSpec, a training-free compiler that derives complementary evidence specifications while preserving the original question for final reasoning. We further validate it through matched-control experiments that isolate the roles of evidence specification and localization. With the search budget fixed, structured evidence specifications yield an 8.6% relative gain over generic requests. With evidence geometry matched, the evidence localized by EviSpec yields a 14.8% relative gain over random evidence. Together, these controls isolate the benefit of specifying what evidence to seek rather than merely expanding visual access. Across all five MLLMs, EviSpec consistently improves upon the corresponding baseline on each of the three benchmarks, yielding average relative gains of \textbf{10.4%, 8.8%, and 12.4%} on V\textsuperscript{*}Bench, HR-Bench-4K, and HR-Bench-8K, respectively. Beyond high-resolution reasoning, EviSpec also achieves state-of-the-art performance on VQA and hallucination-focused benchmarks.
arXiv ID: 2609.23495 / 要約の誤りについて