公開データとモデル規模に応じた学習で画像推論を改善する
MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning
この論文をやさしく読む
ひとことで言うと
画像を含む推論モデルの学習データを整理し、得意分野別に学習したモデルの能力を統合する方法です。
何に役立つ?
公開データで事後学習を設計する際に、データ難易度や選別方法をモデルの規模に合わせる材料になります。
この研究の面白いところ
小さいモデルと大きいモデルで教師データの選別効果が異なり、領域を単純に混ぜるより専門化後の蒸留が安定すると分析しています。
どこまで分かった?
平均72.8、74.4は15ベンチマークの集計値です。個々の分野で一様に改善したとは限らず、実運用での信頼性全般を保証する結果ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
オープンなマルチモーダル推論モデルは、大規模な推論用教師データの恩恵を受けてきた。しかし、データ品質のばらつき、非効率な教師データ構築、難易度の偏り、領域間の干渉により、信頼できる事後学習は依然として難しい。そこで、3つの要素からなる、公開データを使う事後学習手順MMVistaReason(MVR)を導入する。 第1は、構造化された推論を重視する分析的推論群と、視覚知覚・空間的な根拠付けを重視する実世界推論群という、相補的な群をまたいだ幅広い能力のカバーである。第2は、教師あり微調整(SFT)と強化学習(RL)のデータを効率よく構築することである。段階的なクリーニングと注釈付けで異質な公開データを標準化し、難易度を考慮した段階的な教師蒸留と、回答尤度に基づく推論軌跡の選択を組み合わせてMVR-SFT-528Kを構築する。また、MVR-RL-63Kには、モデル規模ごとのフロンティア・フィルタリングを適用する。第3は、専門化してから統合する学習である。相補的なRL専門モデルを学習し、複数教師によるオンポリシー蒸留(MOPD)で能力を統合する。 分析により、教師データの難易度、推論軌跡の品質、モデル容量の間に、容量依存の相互作用があることが分かった。小さい生徒モデルは選別した教師データからより大きな恩恵を受ける一方、大きい生徒モデルは軌跡の変動や混合領域の干渉に頑健である。混合領域のRLはベンチマーク単位で負の転移を生むが、MOPDは安定した能力統合を実現し、望ましいKLの方向はモデル規模によって変わる。15のマルチモーダル・ベンチマークで、MVR-4Bは平均72.8を達成し、Qwen3.5-9B(Instruct)とMMFineReason-8Bを上回る。しかも、MMFineReasonより約70%少ないサンプルを使う。9Bへ拡大すると平均は74.4となり、Qwen3.5-35B-A3B(Instruct)を上回る。全体として、体系的な公開データ構築と容量を考慮した事後学習が、信頼できるマルチモーダル推論に向けた実用的で拡張可能な方法となることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven data quality, inefficient supervision construction, imbalanced difficulty, and cross-domain interference. We introduce MMVistaReason (MVR), an open-data post-training recipe with three components: (1) broader capability coverage across complementary Analytical and Real-World reasoning groups, emphasizing structured reasoning versus visual perception and spatial grounding; (2) efficient SFT and RL data construction, standardizing heterogeneous open data through staged cleaning and annotation, combining difficulty-aware cascaded teacher distillation with answer-likelihood-based trajectory selection to construct MVR-SFT-528K, and applying scale-specific frontier filtering for MVR-RL-63K; and (3) specialize-then-integrate training, which trains complementary RL experts and consolidates their capabilities through multi-teacher on-policy distillation (MOPD). Our analyses reveal a capacity-dependent interaction between supervision difficulty, trajectory quality, and model capacity: smaller students benefit more from selected supervision, while larger students are robust to trajectory variation and mixed-domain interference. Mixed-domain RL introduces benchmark-level negative transfer, whereas MOPD provides consistent capability integration, with the preferred KL direction varying across model scales. Across 15 multimodal benchmarks, MVR-4B achieves an average score of 72.8, outperforming Qwen3.5-9B (Instruct) and MMFineReason-8B while using about 70% fewer samples than MMFineReason. Scaling to 9B improves the average to 74.4, surpassing Qwen3.5-35B-A3B (Instruct). Overall, MMVistaReason demonstrates that systematic open-data construction and capacity-aware post-training provide a practical and scalable path toward reliable multimodal reasoning.
arXiv ID: 2610.01352 / 要約の誤りについて