教師モデルへの問い合わせ予算を抑えるマルチモーダル自己蒸留
BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception
この論文をやさしく読む
ひとことで言うと
画像の細部を学ぶ自己蒸留で、教師モデルに問い合わせる訓練例を選び、問い合わせ数を節約する方法。
何に役立つ?
考えられる用途は、マルチモーダルモデルの訓練費用を抑えながら細かな視覚認識を改善することである。推論時は画像全体を一回で処理する。
この研究の面白いところ
学生モデルのバッチ生成を維持し、教師の情報が役立ちそうな標本だけ選ぶ。選択器の計算に学生の追加の順伝播を要しない。
どこまで分かった?
要旨はベンチマークでの『高い性能』と費用の大幅削減を述べるが、具体的な精度値や削減率を記載していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
マルチモーダル大規模言語モデルは画像全体を処理するとき、重要な根拠が局所的な領域だけにある場合、細かな視覚認識に苦戦することが多い。方策上の自己蒸留(OPD)では、情報量の多い視点から得た特別な視覚知識を、画像全体を使う方策へ移せるが、すべてのロールアウトで教師に問い合わせると教師信号のコストが大きい。本研究は、限られた問い合わせ予算内で教師信号を配分する、予算対応の選択的OPDの枠組みBAS-OPDを提案する。すべてのロールアウトで問い合わせる代わりに、学生モデルによるバッチ全体の生成は維持しつつ、情報量の多い標本を選ぶ。無作為選択、不確かさに基づく選択、学習した有用度に基づく選択を検討する。学習した選択器は、切り離したロールアウト統計と、学生と教師の一致度および教師の確信度から得るオンラインの有用度信号を用い、学生モデルの追加の順伝播なしに問い合わせ価値を推定する。BAS-OPDが変えるのは学習時の教師信号の配分だけで、推論時には画像全体を一回で処理する。細粒度のマルチモーダル知覚ベンチマークの実験では、教師信号のコストを大きく減らしながら高い性能を達成し、予算が限られる場合に選択的OPDが有効であることを示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Multimodal large language models (MLLMs) often struggle with fine-grained visual perception when processing complete images, as critical evidence may only appear in local regions. On-policy self-distillation (OPD) enables transferring privileged visual knowledge from informative views to full-image policies, but querying the teacher for every rollout introduces substantial supervision costs. In this work, we propose BAS-OPD, a budget-aware selective OPD framework that allocates teacher supervision under limited query budgets. Instead of querying all rollouts, BAS-OPD selects informative samples while maintaining full-batch student generation. We explore random, uncertainty-based, and learned utility-based selection strategies, where the learned selector estimates query value from detached rollout statistics and online utility signals derived from student--teacher agreement and teacher confidence without additional student forward passes. BAS-OPD only changes training-time supervision allocation and preserves single-pass full-image inference. Experiments on fine-grained multimodal perception benchmarks demonstrate that BAS-OPD achieves strong performance while substantially reducing teacher supervision costs, highlighting the effectiveness of selective OPD under constrained budgets.
arXiv ID: 2609.25891 / 要約の誤りについて