出力分布の変化を抑える言語モデルの推論時制御
Minimally Invasive Steering of Language Models
この論文をやさしく読む
ひとことで言うと
言語モデルの内部状態を推論時に変えつつ、元の出力分布からのずれを抑える方法。
何に役立つ?
モデルの再学習なしで、選好やコード生成の報酬へ出力を調整する研究に役立つ。要旨の実証は約10億~140億パラメータのモデルでの課題評価。
この研究の面白いところ
分布の変化を測るKLに基づいて介入を抑え、七設定中六設定で平均報酬が最高だった。勾配の近似関係も理論的に示す。
どこまで分かった?
後続部分の項が二次であるという結果は生成長が固定された場合。性能結果は示された七つのモデル・課題設定に限られる。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ロジット計算前のステアリングは、最終隠れ状態にベクトルを加え、固定された言語モデルを推論時の報酬へ適応させる。制約を設けずに報酬を最適化すると、出力分布が大きく変わり、生成品質が落ちる可能性がある。本研究は、生成されるトークン分布の局所的なKL幾何を用いて介入にペナルティーを課す、最小侵襲ステアリングベクトル最適化(MISVO)を提案する。得られるフィッシャーの二次形式は分布の感度を測り、固定された言語モデルの出力層との行列・ベクトル積によって計算できる解析的な勾配を持つ。系列全体のKL勾配を、解析的なフィッシャー項と、後続部分に関するスコア関数項へ厳密に分解する。生成する長さが固定されている場合、後続部分の項はステアリングの大きさについて二次であり、三つのフィッシャー近似が完全なKL勾配と一次まで一致することを示す。MISVOは固定した参照モデルによる近似を用い、モデルのパラメータを更新せず、位置ごとに異なる介入を最適化する。約10億~140億パラメータのモデルを使った選好課題とコード生成課題では、七つのモデル・課題の組み合わせのうち六つで平均報酬が最も高く、多様性と一貫性のスコアはBest-of-Nに近かった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Pre-logit steering adapts a frozen language model to a test-time reward by adding vectors to its final hidden states. Unregularized reward optimization can substantially alter the output distribution and degrade generation quality. We propose Minimally Invasive Steering Vector Optimization (MISVO), which penalizes interventions using the local KL geometry of the induced token distribution. The resulting Fisher quadratic measures distributional sensitivity and admits an analytic gradient computed through matrix--vector products with the frozen language-model head. We derive an exact decomposition of the sequence-level KL gradient into an analytic Fisher term and a suffix score-function term. For a fixed generation horizon, we show that the suffix term is second order in the steering magnitude and that three Fisher surrogates agree with the full KL gradient to first order. MISVO uses the frozen-reference surrogate to optimize position-specific interventions without updating model parameters. Across preference and code-generation tasks on models with approximately 1B--14B parameters, MISVO achieves the highest mean reward in six of seven model--task settings, with diversity and coherence scores close to those of Best-of-N.
arXiv ID: 2609.30218 / 要約の誤りについて