自己蒸留の教師モデルに与える情報量を比較
What Should a Self-Teacher See? Privileged Context Design for On-Policy Self-Distillation
この論文をやさしく読む
ひとことで言うと
自己蒸留で教師モデルに完全な解答を与えるより、解法の方針や問題分類などを与えたほうがよい場合があるかを比べた。
何に役立つ?
数学問題向けの自己蒸留で、教師に与えるヒントの形式と保存コストを設計する参考になる。
この研究の面白いところ
中間的な文脈は4B・8Bモデルで完全解答より最高平均得点が1.4・1.6点高く、ヒントのトークン数は一桁少なかった。一方、答えだけの条件も完全解答に近い成績だった。
どこまで分かった?
主な結果は競技数学の4B・8Bモデルを用いた実験に基づく。最適な文脈はモデル規模や課題で変わると報告されており、単一の形式が常に最良とは示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
特権的な情報を多く与えれば、必ずしも良い教師になるわけではない。本研究は、方策上の自己蒸留(OPSD)でこの問題を調べる。OPSDでは、元のモデルを固定したコピーが、学生モデル自身の生成過程を特権的な文脈の下で採点する。通常その文脈には、最終答えと特定の推論経路を一緒に含む完全な模範解答を使う。各モデル規模で、学生モデルが見える情報と学習条件を固定し、この標準設定を、事前に作成した三つの抽象化、すなわち名称のある戦略、特定の解法によらない枠組み、問題の分類と比較した。さらに、答えという到達点だけを残して経路を除いた対照条件も設けた。競技数学を用いた主な実験では、最良の中間的な文脈は、完全な解答に比べて、同一分野の最高平均得点を4Bモデルで1.4点、8Bモデルで1.6点上げ、保存するヒントのトークン数は一桁少なかった。乱数種三つにわたる比較でも、解法によらない枠組みと問題の分類は、両規模で平均値が向上した。答えだけを与える条件も主な実験では競争力があり、これらの規模で完全な解答との差は0.2点以内だった。望ましい文脈は学生モデルの規模と作業によって変わり、初期の教師・学生間のKL距離の大小は、後の性能の順序と一致しなかった。自己蒸留の教師に見せるべきなのは、与えられるすべての情報ではなく、学生が利用できる抽象度の情報である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
More privileged information does not always make a better teacher. We study this tension in on-policy self-distillation (OPSD), where a frozen copy of the base model scores the student's own rollouts under privileged context, conventionally a complete reference solution that bundles the final answer with one particular reasoning path. Holding the student view and training fixed within each scale, we compare that default against three abstractions compiled offline, a named strategy, a method-independent framing, and a problem category, and against an answer-only control that keeps the destination but removes the path. In the primary runs on competition mathematics, the best intermediate contexts improve the in-domain peak mean over the full solution by 1.4 points at 4B and 1.6 at 8B, while storing an order of magnitude fewer hint tokens. Comparisons across three seeds also show positive mean gains for the framing and category contexts at both scales. Answer-only conditioning remains competitive in the primary runs, within 0.2 points of the full solution at these scales. The preferred context varies with student scale and task. Initial teacher-student KL does not order downstream performance. What a self-teacher should see is therefore not everything it could, but the level of abstraction its student can still act on.
arXiv ID: 2609.25623 / 要約の誤りについて