arXiv論文メモ
新着一覧
cs.CL / cs.AI / cs.LG · 査読状況未確認

比較結果から言語モデルを好みに合わせるComPO

A Zeroth-Order Paradigm for LLM Preference Alignment

Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin

この論文をやさしく読む

ひとことで言うと

回答ペアの好ましさの比較から更新方向を取り出し、言語モデルを人間の好みに近づける手法です。

何に役立つ?

選好ペアの尤度差が小さい状況でも比較情報を活用する、モデル調整の選択肢になります。複数のモデル系列で既存手法と比較されています。

この研究の面白いところ

ペアに対する微分可能な損失を直接最適化せず、比較オラクルを使います。オフライン方式と、参照方策から離れすぎないよう制御するオンライン方式を扱っています。

どこまで分かった?

理論保証には滑らかさ、勾配の疎性、オラクルの整合性などの仮定があります。尤度変位の緩和は診断と整合する証拠として述べられ、要旨に具体的な改善幅はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

直接的な選好調整手法は、計算とメモリの効率がよいため、大規模言語モデル(LLM)を人間の選好に合わせるために広く使われている。しかし、尤度変位の問題は、尤度の差が小さい選好ペアから情報を抽出する別の方法を検討する動機となる。本論文では、比較オラクルに基づくゼロ次の調整手法である比較ベース選好最適化(ComPO)を提案し、解析する。ComPOは、これらのペア上で微分可能な選好損失を直接最適化せずに、ペアから方向の情報を抽出する。 滑らかさ、勾配の疎性、オラクルと潜在目的関数の整合性という条件の下で、基本的なオフライン方式の収束保証を示す。さらに、オフラインの比較機構を保ちつつ、ラベルのない方策生成文を使って参照方策に対する逆KLを制御する、オンラインComPOを導入する。選好ファインチューニングをカバレッジの観点から捉え、局所的なカバレッジと分布内のペアごとの報酬精度の下で、基本的な制約付き方式の性能保証を確立する。 Mistral、Llama、Gemma-2、Qwen3、Gemma-3モデルでの実験では、長さを調整した勝率を含め、既存の直接調整手法に対する改善が示される。ペア単位の診断からも、尤度変位の緩和と整合する証拠が得られる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margins. In this paper, we propose and analyze Comparison-based Preference Optimization (ComPO), a zeroth-order alignment method based on comparison oracles. ComPO extracts directional information from these pairs without directly optimizing a differentiable preference loss on them. We establish a convergence guarantee for its basic offline scheme under smoothness, gradient sparsity, and compatibility between the oracle and a latent objective. We further introduce online ComPO, which retains the offline comparison mechanism and uses unlabeled policy generations for reverse-KL control relative to a reference policy. Following the coverage perspective of preference fine-tuning, we establish a performance guarantee for a basic constrained scheme under local coverage and in-distribution pairwise reward accuracy. Experiments on Mistral, Llama, Gemma-2, Qwen3, and Gemma-3 models demonstrate improvements over existing direct alignment methods, including length-controlled win rates, with pair-level diagnostics providing evidence consistent with mitigating likelihood displacement.

著者のコメント

39 pages

arXiv ID: 2609.19144 / 要約の誤りについて