arXiv論文メモ
新着一覧
cs.CL / cs.IT / math.IT · 査読状況未確認

履歴を持つ言語モデルで二つの調整法の分布を比較する

The Asymptotics of Language Model Alignment with Memory

Haricharan Balasundaram, V. Arvind Rameshwar

この論文をやさしく読む

ひとことで言うと

報酬を高めるようモデルを学習し直す方法と、複数の生成候補から良いものを選ぶ方法が、どの条件で似た出力分布になるかを調べています。

何に役立つ?

アラインメント手法を分布の観点から比較し、サンプリングによる選別とKL制約付き強化学習の関係を理解するために役立ちます。

この研究の面白いところ

トークンが互いに独立という仮定を、直前までの状態に依存するマルコフ的な列へ広げています。長い列で近づく話だけでなく、有限長、特に1トークンで完全に一致する条件も扱います。

どこまで分かった?

マルコフ的な出力列に関する理論結果で、任意の履歴依存を持つ実用LM全般で同じ結論が成り立つとは述べていません。要旨には有限の長さでの数値的な誤差や実モデルでの性能比較はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

言語モデル(LM)のアラインメントは大まかには、与えられたLMのQを調整済みLMのqへ変化させ、(i)qとQが生成する出力が確率的に「近い」こと、(ii)qの期待報酬がQより高いことを目指す。よく使われる二つの手法は、LMの分布の知識が必要で計算費用が大きいKL制約付き強化学習と、LMからのサンプリングだけを必要とするbest-of-nアルゴリズムである。 Yangらは、LMが出力する長さmのトークン列が独立同分布である場合、mを無限大にする極限で、二つの手法が生成する分布が漸近的に近くなることを示した。しかし、実際のLMの出力列はしばしば記憶を持つため、独立同分布の仮定は実用的なLMを代表していない。本論文では、長さmの出力トークン列がマルコフ的である場合へ、この漸近的近接性の結果を拡張する。さらに、有限長の出力列、特にm=1の場合について、二つの手法が生成する分布間のKLダイバージェンスがゼロとなるLM分布と報酬関数を完全に特徴付ける。これはYangらが最初に提起した問題である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowledge of the LM distribution and is computationally expensive, and the best-of-$n$ algorithm, which requires only sampling from the LM. The work of Yang et al. established asymptotic closeness between the distributions produced by the two alignment methods for an $m$--length i.i.d. token sequence output by the LM, in the limit as $m$ increases to infinity. However, the i.i.d. assumption is not representative of practical LMs, whose output sequences often have memory. In this paper, we extend the asymptotic closeness result to the case when the $m$--length token sequence outputted by the LM is Markovian. Further, for finite-length output sequences -- particularly, when $m=1$ -- we provide a complete characterization of LM distributions and reward functions for which the KL-divergence between the distributions produced by the two alignment methods is zero -- a question first posed in Yang et al.

arXiv ID: 2610.01828 / 要約の誤りについて