arXiv論文メモ
新着一覧
cs.LG / cs.AI / cs.CL · 査読状況未確認

LLM強化学習で確率変化の打ち消しを防ぐ応答選別

CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning

Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang

この論文をやさしく読む

ひとことで言うと

LLMの強化学習で古い方策の応答を選別するとき、トークン確率の増減が平均で相殺される問題を防ぐ方法です。

何に役立つ?

生成と学習で方策がずれやすい環境で、学習に使う応答の選択を改善する用途が考えられます。

この研究の面白いところ

対数比の絶対値を先に取る小さな変更を、範囲外の割合と逸脱量の理論保証につなげています。

どこまで分かった?

報告値は記載された数学・コードベンチマークでの比較です。mean@16とpass@1は異なる指標であり、最大3.13ポイントという値を全課題共通の改善量と読むことはできません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

近年、大規模言語モデル(LLM)の事後学習で強化学習(RL)の利用が急速に広がり、数学的推論やコード生成が大きく改善してきた。しかし実際のシステムでは、方策更新と、応答生成用・学習用エンジンの違いによって、標本化した応答が現在の方策と異なるオフポリシーになることがある。系列単位のマスキングは、応答全体を最適化へ使うかどうかを決めることで、この不一致に対処する。よく使われる規則は、標本化したトークンの確率比の、長さで正規化された幾何平均を用いる。しかし、符号付き対数比は位置間で打ち消し合い、双方向の大きな方策のずれを隠し得る。 そこで、平均する前に各トークンの対数比の絶対値を取る系列単位のマスク、Cancellation-Aware Response Masking(CARM)を提案する。これにより、互いに逆向きの確率変化が打ち消し合うことを防ぐ。採用された応答について、所定の範囲外にある標本トークン比の割合と、その境界を超えた平均対数距離に同時の上界が成立することを証明する。 数学的推論とコード生成の実験では、CARMはAIME 2024・2025・2026およびBeyondAIMEで平均したmean@16を、幾何平均マスキングより最大3.13パーセントポイント改善する。また、四つのコードベンチマークの平均pass@1を、評価した最も強い比較手法より2.88ポイント高める。これらの知見は、CARMがLLM強化学習の応答単位のオフポリシー制御に対して、理論的根拠を持つ有効な方法であることを裏付ける。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.

著者のコメント

28 pages, 11 figures, 5 tables

arXiv ID: 2610.02039 / 要約の誤りについて