画像と質問の変換で視覚言語モデルの空間推論を改善する
INTCORT: Training-Free Spatial Reasoning Enhancement for Vision-Language Models via Input Transformations and Confidence Routing
この論文をやさしく読む
ひとことで言うと
画像の向きや質問の聞き方を変えて複数回答を得て、「左」「右」などの関係を表す語の信頼度を使って回答をまとめる方法です。
何に役立つ?
モデルを追加学習したり内部構造を変えたりせずに、位置関係を問う課題の精度を改善する用途が考えられます。
この研究の面白いところ
入力を変えると誤答が正答に変わることと、関係トークンの信頼度が正誤と対応することを組み合わせています。複数の見方を作り、信頼度に基づいて集約します。
どこまで分かった?
要旨は平均10.01%、先行研究比で最大25.01%の改善と表記していますが、相対改善率かパーセントポイント差かは明記していません。追加推論のコストや個々のモデルの成績も要旨にはありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
視覚言語モデル(VLM)はマルチモーダルな課題で優れた能力を示してきたが、空間推論の能力は依然として低い。既存の改善法のうち、学習を伴う方法は高い計算コストと破滅的忘却に、学習不要の方法は一般的な能力を損なう内部機構への干渉に、それぞれ課題を抱えている。本研究ではまず、適切な幾何学的画像変換と質問の反転変換によって誤った空間予測を正せること、正しい予測は誤った予測より関係を表すトークンの信頼度が高いこと、という2つの重要な仮説を検証する。 この知見に基づき、VLM内部の機構を変更せず、入力変換で複数の推論用ビューを作り、関係トークンの信頼度による振り分けを通じて予測を集約する、学習不要の空間推論改善フレームワークINTCORTを提案する。広く使われる複数のベンチマークでの実験は、多様なVLMの空間推論精度が大幅に改善することを示し、全モデルと全ベンチマークを通じた平均改善は10.01%となった。先行研究との比較では、最大25.01%の改善を伴う優れた性能を達成した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vision-Language Models (VLMs) have demonstrated remarkable capabilities in multimodal tasks, yet they still exhibit poor ability in spatial reasoning. Existing training-dependent and training-free enhancement methods suffer from high computational costs with catastrophic forgetting and internal mechanism interference that compromises general capabilities, respectively. In this work, we first verify two key hypotheses: appropriate geometric image transformation and query-reversal transformation can recover incorrect spatial predictions, and correct predictions exhibit higher relation-token confidence than incorrect ones. Based on these findings, we propose INTCORT, a training-free spatial reasoning enhancement framework that constructs multiple inference views through input transformations and aggregates their predictions via relation-token confidence routing, without modifying the VLM's internal mechanisms. Experimental results on several commonly-used benchmarks demonstrate that INTCORT substantially improves spatial reasoning accuracy across diverse VLMs, achieving an average improvement of 10.01% over all models and benchmarks. Compared with prior works, INTCORT achieves superior performance with improvements of up to 25.01%.
arXiv ID: 2609.24813 / 要約の誤りについて