画像言語モデルの規模と画像解像度の配分を調べる
Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
この論文をやさしく読む
ひとことで言うと
同じ計算予算なら言語モデルを大きくするべきか、画像を詳しく見せるべきかを、測定結果に基づく法則で選ぶ研究です。
何に役立つ?
高解像度画像を扱うVLMの運用で、利用可能なモデルと画像サイズから予算に合う組合せを決める際に役立ちます。
この研究の面白いところ
モデルを大きくしても改善しない質問が相当数あり、言語部分の拡大と視覚情報の増加では効き方も違うと示しています。
どこまで分かった?
法則は二つのモデル系列の26モデルと四つのベンチマークの測定に当てはめたものです。要旨は、あらゆるVLMや課題への普遍的保証を示しているわけではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
画像言語モデル(VLM)には、一定の予算の下で、細部を知覚するためにより多くの視覚情報を処理することと、複雑な推論のためにより大きな言語バックボーンを使うことの間にトレードオフがある。既存研究からは、特に高解像度での運用において、どのバックボーン規模と入力解像度の組合せを採用すべきかが分からない。 この不足に対処するため、言語バックボーンの規模と視覚トークン数によってVLMの性能がどう変わるかを記述する「分離可能則(Separable Law)」を提案する。言語バックボーンのパラメータ数が10億から720億のInternVLおよびQwenVLの26モデルを対象に、画像サイズが224ピクセルから8Kに及ぶ四つの高解像度ベンチマークで得た測定値へ、この法則を当てはめる。規模拡大に反応する質問は、その質問が必要とする能力から予測できる一方、かなりの割合の質問は規模拡大にまったく反応しないことが分かった。また、二つのモデル系列はバックボーンを大きくすることで同程度の改善を得るが、視覚トークンを増やしたときの改善には大きな違いがある。 分離可能則をコスト則と組み合わせると、バックボーン規模と視覚トークンの間で計算量を配分する閉じた形の規則が得られる。運用時に利用可能な構成が限定される場合でも、この法則は、同じ予算内で実現可能な最良の選択に近い性能を示すモデル規模と画像サイズを特定する。本研究が、モデルに求められる推論内容に応じて、どれだけ高解像度の視覚情報を与えるべきかを判断するための、原理に基づいた方法になることを期待する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they require, while a substantial fraction never responds at all. We also find that the two model families gain similarly from a larger backbone, while their gains from more visual tokens differ sharply. Combined with a cost law, the Separable Law gives a closed-form rule for allocating compute between backbone size and visual tokens. When deployment is limited to available configurations, the law identifies model and image sizes that perform close to the best feasible choice under the same budget. We hope our work offers a principled way to decide how much a model should be allowed to see at high resolution, given what it must reason about.
arXiv ID: 2610.01640 / 要約の誤りについて