arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

長い推論で報酬の差が消える問題を抑えるLLMの強化学習

Rethinking Probability-Based Reinforcement Learning From Posterior Concentration

Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao

この論文をやさしく読む

ひとことで言うと

推論を長く書くほど解答の良し悪しを報酬で区別しにくくなる問題を調べ、学習用データの選び方と長さへの罰則で対処します。

何に役立つ?

外部の正誤判定器を用意しにくい推論タスクで、LLMの強化学習を効率化する用途が考えられます。

この研究の面白いところ

長い推論なら常に悪いとするのではなく、参照解答の事後確率の差が小さくなり、学習信号が弱まる状況を捉えて対策しています。

どこまで分かった?

7つ中6つのベンチマークで改善を報告しており、すべてで優位だったわけではありません。最大4.0%という改善が相対比かパーセントポイント差かは、要旨では明記されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

確率に基づく報酬を使う検証器不要の強化学習は、外部の検証器を用意できない一般的な推論タスクでLLMを学習させる有望な方法である。しかし、特に長い推論における報酬の信頼性は十分に調べられていない。本研究は、確率報酬の長さに依存する失敗の様式を特定し、事後確率集中現象(PCP)と呼ぶ。推論過程を条件とする参照解答の確率は、その過程が長くなるにつれ、分散の小さな区間へ収縮することが多いと示す。この現象により報酬同士がほとんど区別できなくなり、GRPOに基づく設定では、確率を用いた方策最適化が不安定かつ非効率になる。 これを踏まえて、PCPを明示的に考慮し、最適化の安定性とトークン効率を高める、検証器不要の強化学習の枠組みRLCPR(Reinforcement Learning with Concentration-aware Posterior Rewards)を提案する。構成要素は2つある。1つは、不確実性を考慮したデータ抽出で、生成前に確率の集中が起こりやすいロールアウトを減らす。もう1つは、集中を考慮した正則化で、事後確率報酬が潰れてしまった際に不必要に長い推論過程へ罰則を与える。広範な実験では、RLCPRはトークン効率を高めるとともに、一般分野と数学の推論課題を含む7つのベンチマークのうち6つで、最先端の検証器不要の強化学習ベースラインを最大4.0%上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.

arXiv ID: 2610.01458 / 要約の誤りについて