画像が正解に不可欠な課題で視覚的な探索能力を学ぶ
Same Reward, Different Skills: When Multimodal RL Learns to Look
この論文をやさしく読む
ひとことで言うと
視覚言語モデルの点数が上がっても、画像を見る能力が育ったとは限りません。画像を見ないと解けない課題にすると、対象を探す能力が向上したという研究です。
何に役立つ?
視覚モデルの学習課題や評価を設計する際、文章だけで正解できる抜け道がないかを点検する手掛かりになります。
この研究の面白いところ
画像を灰色に置き換える対照実験と、答えを含むキャプションを加える実験で、報酬の同じ設定でも何を学ぶかが変わることを確かめています。
どこまで分かった?
報告された数値は特定のモデル規模、座標場面、対応付け課題での実験結果です。あらゆる画像理解課題で同じ改善率になることまでは示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
検証可能な報酬を用いる強化学習(RLVR)は、学習中に視覚情報がなくても視覚言語ベンチマークのスコアを改善する。テスト時に画像を与えると、画像を見ずに学習したモデルは、実画像を用いた学習による改善のうち、30億パラメータ規模ではおよそ半分、70億規模ではほぼ5分の4を達成する。実画像での学習を長く続けると、ベンチマークの改善が持続していても、画像との対応付けが損なわれる場合がある。両方の結果が示す隔たりは同じである。プロンプトに画像があることは、学習信号に画像が含まれることを意味しない。 私たちの設計原則である「視覚的に解決可能であること」は、正答に視覚的証拠が必要であり、なおかつ課題が学習可能であることを求める。この原則を、質問を固定し、対象の名前を一切示さない反実仮想的な座標場面で検証する。このため、正解には画像中の対象を見つける必要がある。標準的なGRPOと正確性・形式の報酬を用いると、70億パラメータのモデルは、学習時のどの場面よりも密度の高い未使用の場面において、対象を見つける精度(発見精度)を0.425から0.875へ高め、学習していない種類の質問でも改善する。 2つの対照実験によって改善の由来を特定する。テスト画像を灰色の画像に置き換えると、発見精度はゼロになる。一方、灰色の画像で学習した場合、学習ステップを30にそろえた比較では、4つの乱数シードのいずれでも、その後のテストに実画像を使っても改善はほぼ得られない。学んだ技能は、学習コーパスとは独立に構成した対応付け課題にも移る。同じ画像・報酬・予算のまま、学習時の質問への答えとなるキャプションを加えると、改善量はほぼ3分の2減少する。報酬を得るために必要なものを変えると、強化学習が学ぶものも変わる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the image. With standard GRPO and correctness-and-format rewards, a 7B model raises its accuracy at finding the target (discovery) from 0.425 to 0.875 on held-out scenes denser than any it trained on, and it improves on question types it never trained on. Two controls locate the source of the gain. Replacing test images with gray canvases drops discovery to zero; training on gray canvases instead, at matched step 30 and in each of four seeds, yields essentially none of the gain even when the model is then tested with real images. The learned skill carries over to grounding tasks built independently of the training corpus. A caption that answers the training question, added to the same images, reward and budget, cuts the gain by nearly two thirds. Changing what reward requires changes what RL learns.
arXiv ID: 2610.01908 / 要約の誤りについて