arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

自己蒸留で正しさと回答の振る舞いを切り分ける

On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation

Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub, Frederike Lübeck, Jonas Hübotter, Thomas Kleine Buening, Andreas Krause

この論文をやさしく読む

ひとことで言うと

正解を知る教師の回答をまねると、正しさだけでなく回答の長さや自信の出し方も変わります。正解の教師へ近づけ、誤答の教師から遠ざける組合せで、その混同を減らす研究です。

何に役立つ?

推論能力を改善しながら、回答が過度に短くなったり長くなったりする変化を抑える学習設計に役立ちます。ここで示す成果はモデルの推論性能と回答長に関するものです。

この研究の面白いところ

正解教師の模倣だけでなく、誤答教師から離れる信号を加えると、両教師に共通する振る舞いの変化を相殺できるという考え方です。GRPOと切り離して自己蒸留自体の効果を調べています。

どこまで分かった?

要旨は複数のモデル種別での改善を述べていますが、個々のベンチマーク値や改善率は記載していません。あらゆるモデルや課題で同じ安定性が得られるとまでは読み取れません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

オンポリシー自己蒸留は、モデルに特権的な情報を条件として与え、得られた教師分布をそのモデル自身へ蒸留することで、密なトークン単位の教師信号を提供する。しかし特権的な情報は、教師が知っている内容だけでなく、その振る舞いも変え得るため、正しさに関係する学習信号と意図しない振る舞いの変化が絡み合う。 本研究では推論課題を用い、モデルを特権情報を持つ教師に近づける引力的自己蒸留と、その教師から遠ざける斥力的自己蒸留を対比してこの効果を調べる。両目的は、強く、互いに逆向きの振る舞いの変化を誘発し得ることが分かった。引力は探索的な推論を抑え、短く自信の強い回答を促す。一方、斥力は回答を長くし、モデルが潜在的に持つ思考モードへの意図しない切替えを引き起こすことがあり、最終的には不安定になる。 この観察を受け、正しい解を条件とする教師への引力と、誤った解を条件とする教師からの斥力を組み合わせる対比的自己蒸留を調べる。これらの蒸留信号をGRPO目的と組み合わせる先行研究とは異なり、自己蒸留の目的を切り離し、単独の振る舞いを検討する。二つの教師に共通する振る舞いの変化は大部分が打ち消し合い、正しさをより直接的に反映するトークン単位の信号が残ることが分かった。非思考モデル、指示応答専用モデル、すでに思考するモデルのいずれでも、この対比的目的は回答長を安定して保ちながら推論性能を改善する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

On-policy self-distillation provides dense, token-level supervision by conditioning a model on privileged information and distilling the resulting teacher distribution back into the model. However, privileged information can change not only what the teacher knows, but also how it behaves, entangling correctness-relevant learning signals with unintended behavioral shifts. We study this effect in reasoning tasks by contrasting attractive self-distillation, which moves the model toward a privileged teacher, with repulsive self-distillation, which moves it away from a privileged teacher. We find that both objectives can induce strong and opposing behavioral shifts: attraction suppresses exploratory reasoning and promotes shorter, more confident responses, whereas repulsion increases response length, can trigger unintended switches into a model's latent thinking mode, and ultimately becomes unstable. Motivated by these observations, we study contrastive self-distillation, which combines attraction toward a correct-solution-conditioned teacher with repulsion from an incorrect-solution-conditioned teacher. In contrast to prior work that combines such distillation signals with a GRPO objective, we isolate the self-distillation objective and study its behavior on its own. We find that the shared behavioral shifts of the two teachers largely cancel, leaving a token-level signal that more directly reflects correctness. Across non-thinking, instruct-only, and already-thinking models, this contrastive objective improves reasoning performance while maintaining stable response lengths.

arXiv ID: 2609.21561 / 要約の誤りについて