少ない利用者評価でLLMを継続的に個人化するCOPE
COPE: Continual Personalization of LLMs under Sparse User Feedback via User Embeddings and Self-Evaluation
この論文をやさしく読む
ひとことで言うと
利用者の評価が少なくても、利用者ごとの好みを学びながら応答を更新するLLMの手法です。
何に役立つ?
利用者ごとに応答を調整する仕組みの研究に役立ちます。要旨では比較手法に対する実験結果を報告しています。
この研究の面白いところ
明示的な評価がないときにも自己評価から代理報酬を作り、個人化と継続更新を一回の更新にまとめています。
どこまで分かった?
要旨は具体的な評価値や実利用環境での長期的な成果を示していません。自己評価の信頼性は論文の実験条件下で確認されたものです。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルはさまざまなベンチマークで優れた結果を出しているが、規範的価値観への適合によって応答が均質化し、多様な利用者の好みに対応できないことがある。学習を伴わない既存手法はプロンプトの工夫に貴重な文脈窓を使うことが多く、学習を伴う手法は学習後に固定されがちで、現実の利用で必要な継続的な最適化を支えられない。そこで、利用者からの評価が少ない実際の対話状況を想定した最適化の枠組みCOPEを提案する。利用者ごとに学習可能な個人化埋め込みを割り当て、好みの把握、自己評価の較正、個人化した応答の最適化を一回の更新に統合する。重要な工夫は、自己評価で代理報酬を生成し、明示的な利用者評価がない場合でもモデルを継続して更新できるようにする点である。実験では、評価が少ない条件で、学習しない手法と学習する手法の強力な比較対象を一貫して上回り、検索拡張プロンプティングとも補完的に機能した。追加分析では、自己評価の信頼性、意味のある好みのパターン、一般能力の安定性、好みの変化や別の評価器に対する頑健性も確認した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
While Large Language Models (LLMs) have achieved remarkable results across various benchmarks, their alignment with normative values often results in homogenized responses that fail to address diverse user preferences. Existing training-free methods often occupy valuable context windows through prompt engineering, while training-based methods typically remain static post-training, failing to support the continual optimization required in real-world settings. To address these challenges, we propose COPE (Continual Optimization with Personalized embedding and self-Evaluation), a novel optimization framework tailored for real-world-motivated interaction settings with sparse user feedback. Our framework assigns learnable personalized embeddings to each user and synergistically integrates preference capture, self-evaluation calibration, and personalized response optimization within a single update step. A key innovation of our method is the use of self-evaluation to generate proxy rewards, enabling continuous model updates even when explicit user feedback is unavailable. Experiments show that COPE consistently outperforms strong training-free and training-based baselines under sparse feedback, and remains complementary to Retrieval-Augmented Prompting (RAP). Further analyses confirm COPE's reliable self-evaluation, meaningful preference patterns, stable general capabilities, and robustness under shifting preferences and alternative evaluators.
arXiv ID: 2609.26853 / 要約の誤りについて