言語モデルの評価点だけを最適化する危険を三つの方法で比較
Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
この論文をやさしく読む
ひとことで言うと
言語モデルの評価点を上げる操作が、実際の課題性能を改善せずに採点の穴を突く場合を比較した理論的研究。
何に役立つ?
モデル更新、出力選択、固定プロンプトの改善を評価するとき、点数とは独立の品質確認を設計する助けになる。
この研究の面白いところ
三つの最適化対象を同じ枠組みで比較し、評価器との距離だけでは報酬ハッキングの起きやすさを一律に順位付けできないと示す。
どこまで分かった?
形式解析、有限の数値例、既報の証拠に基づく枠組みであり、新しい大規模な運用実験の結果は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
評価点が上がっても、言語モデルシステムの課題遂行が良くなったとは限らない。最適化が評価器の誤りを利用すると、測定上の進歩の裏で、実際の性能が変わらないか悪化することがある。この失敗は、モデルのパラメータ更新、生成候補からの選択、持続的なプロンプトの書き換えのいずれでも起こり得る。本研究は、この三つの最適化対象で報酬ハッキングを比較する枠組みを作る。代理指標の圧縮仮説と、推論時・文脈内の報酬ハッキングの研究を基に、到達可能な振る舞い、最適化予算、持続的な適応が代理指標の誤差への接触をどう変えるかを調べる。評価器の不一致について距離に依存する上界と、入れ子になった方策クラスの容量の順序を定式化し、距離だけでは脆弱性の普遍的な順位を決められないことを示す。有限個の出力からなる厳密な例では、採点の欠陥がどこにあるかで各方法に有利な振る舞いが変わる。代表的な防御法を三つの対象に対応付け、直接移せるものと機能だけが類似するものを区別する。持続的なプロンプトは中身を調べられる一方、小さな文言変更が誘発する振る舞いは予測しにくいため、特に注意を払う。形式解析、数値例、既報の証拠を合わせ、最適化方法と防御法の適用条件を比較する基礎を示す。信頼できる改善には、到達可能な失敗を管理し、最適化する点数とは独立した課題の質の証拠を保つ必要がある。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator's mistakes, measured progress can conceal unchanged or deteriorating task performance. This failure can arise through parameter updates, selection among generated outputs, or revisions to persistent prompts. We develop a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text. Building on the Proxy Compression Hypothesis and research on inference-time and in-context reward hacking, we examine how reachable behavior, optimization budgets, and persistent adaptation shape exposure to proxy error. We formalize a distance-dependent upper bound on evaluator disagreement and a capacity ordering for nested policy classes, then show why distance alone cannot establish a universal ranking of vulnerability. An exact finite-output illustration demonstrates how the location of a scoring defect changes the behavior favored by each method. We also map representative defenses across substrates, identifying which mechanisms transfer directly and which offer only functional analogies. Persistent prompts receive particular attention: their contents are inspectable, but the behavior induced by a small textual change may be difficult to anticipate. The formal analysis, numerical illustration, and published evidence together provide a basis for comparing optimization methods and identifying the conditions under which their defenses transfer. The resulting framework connects optimization choices to verification requirements: reliable improvement depends on controlling accessible failure modes and preserving evidence of task quality independent of the score being optimized.
著者のコメント
19 pages, 1 figure, 2 tables
arXiv ID: 2609.25848 / 要約の誤りについて