合成画像の位置情報を使って画像理解モデルを自己蒸留する
Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
この論文をやさしく読む
ひとことで言うと
合成画像の中で「どこを見るべきか」を教師側だけに教え、最終的には画像と質問だけで答えるモデルを育てる方法です。
何に役立つ?
人が位置注釈を付ける費用をかけず、画像の数え上げや文書・グラフ理解を追加学習する用途が考えられます。実画像の評価へ転移した結果も報告されています。
この研究の面白いところ
画像の拡大そのものではなく、関連する複数領域の場所をテキストで案内します。合成場面だけで学んだ改善が、複数の実世界ベンチマークにも現れています。
どこまで分かった?
3.23ポイントは列挙された6つの評価の平均改善です。要旨には各モデル・各評価の詳細な得点や、すべての実画像条件への一般化保証はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
オンポリシー自己蒸留は、特別な情報を与えた、固定版または指数移動平均(EMA)版の自分自身を教師として生徒を監督することで、言語モデルの推論を改善する有効な方法として最近登場した。しかし、マルチモーダル大規模言語モデル(MLLM)への適用は、ほとんど研究されていない。最近の方法は、質問に対応する画像の切り抜きなどの特別な視覚情報を使って細かな知覚を改善するが、その改善は視覚的な拡大が有用な課題に限られ、人による位置対応の注釈データか、外部の教師モデルを必要とする。 本研究は、問い合わせに関連する視覚要素を特定する、空間的位置に根差したテキストの案内を教師へ与える、別の形式のMLLM向けオンポリシー自己蒸留を導入する。物体の識別情報と空間座標を自動的に得られる、手続き的に生成した場面を使い、拡張可能で注釈不要の事後学習を実現する。教師はこの空間的な案内を使い、関連する複数の画像領域から証拠を見つけて統合する。一方、生徒は画像と質問だけから、その振る舞いを再現するよう学ぶ。 この方法は、複数のモデルを通じて、数え上げ、文書理解、グラフ理解のベンチマークの性能を一貫して改善する。事後学習には合成場面だけを使うにもかかわらず、改善は実世界の知覚ベンチマークへ転移し、CVBench、V*、ZoomBench、BLINK、HR-Bench、MME-RealWorldの平均性能で3.23ポイントの向上をもたらす。これらの結果は、空間的位置に根差した特別な情報が、オンポリシー自己蒸留を通じてより広い知覚能力を引き出し、事後学習で用いた課題とデータ分布を越えて、合成データから実データへの大きな転移を可能にすることを示す。プロジェクトページ:https://github.com/sirkosophia/Where-OPD
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance identifying the visual elements relevant to a query. We use procedurally generated scenes with automatically available object identities and spatial coordinates, enabling scalable and annotation-free post-training. The teacher uses this spatial guidance to locate and integrate evidence from multiple relevant image regions, while the student learns to reproduce the resulting behavior from the image and question alone. Our approach consistently improves performance on counting, document and chart understanding benchmarks across multiple models. Importantly, although post-training uses only synthetic scenes, the resulting improvements transfer to real-world perception benchmarks, yielding a 3.23-point gain in average performance across CVBench, V*, ZoomBench, BLINK, HR-Bench, and MME-RealWorld. These results show that spatially grounded privileged information can induce broader perceptual capabilities through on-policy self-distillation, enabling substantial synthetic-to-real transfer beyond the task and data distribution used for post-training. Project page: https://github.com/sirkosophia/Where-OPD
arXiv ID: 2610.02117 / 要約の誤りについて