教師自身の変動を除いて推論モデルへ知識を移す
Calibrating Teacher--Student Discrepancy for On-Policy Distillation
この論文をやさしく読む
ひとことで言うと
大きなAIから小さなAIへ教えるとき、両者の出力の違いをすべて学習対象にせず、教師自身が条件によって揺れる部分を取り除く方法です。
何に役立つ?
数学的推論モデルの蒸留で、学習信号の質を改善する用途が考えられます。要旨では複数のモデル規模で改善を報告していますが、数学以外の課題の結果は示していません。
この研究の面白いところ
教師と生徒の違いを能力差と同一視せず、教師に情報を与える介入から教師自身の変動範囲を推定します。元の信号の約52〜65%に絞っても比較手法を上回ったという結果です。
どこまで分かった?
52〜65%は保持した最適化信号の割合であり、計算時間や必要メモリの削減率ではありません。性能改善の具体的な点数や数学以外への一般化については要旨に記載がありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
オンポリシー蒸留(OPD)は、より強力な教師と、現在の方策に従う生徒とのトークン単位の差を学習することで、推論モデルを改善する。しかし、この差が純粋に教師と生徒の能力差だけを反映するわけではない。教師自身に由来するずれも含まれ、それが観測される教師・生徒間の差に混ざり、標準的なOPDは訓練中にそれを区別せず学習する。この問題は、特権情報を用いるOPDでさらに深刻になる。特権情報によって教師側の尤度が大きく変動し、その結果、生徒が教師自身のずれをより多く学ぶよう促されるためである。 本研究ではCalibrated On-Policy Distillation(Cal-OPD)を導入する。正と負の特権的介入を通じて教師自身のずれの領域を推定し、その領域を超える成分だけを残すことで、元の教師・生徒間の差を較正する。数学的推論のベンチマーク実験では、元の教師・生徒間の差の約52〜65%だけを最適化信号として保持しながら、Cal-OPDはモデルの規模をまたいで標準的なOPDとその派生手法を一貫して上回った。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger teacher-side likelihood shifts, thereby encouraging the student to learn more of the teacher's own deviation. We introduce \textbf{Calibrated On-Policy Distillation (Cal-OPD)}, which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region. Experiments on mathematical reasoning benchmarks show that, while retaining only about 52--65\% of the original teacher--student discrepancy as the optimization signal, Cal-OPD consistently outperforms standard OPD and its variants across model scales.
arXiv ID: 2609.21619 / 要約の誤りについて