arXiv論文メモ
新着一覧
cs.LG / cond-mat.dis-nn / stat.ML · 査読状況未確認

検証可能な報酬による学習の最適化地形を解析

RLVR landscapes for iterated multiplications can be benign: Insights from spin-glass theory

Noa Rubin, Zohar Ringel

この論文をやさしく読む

ひとことで言うと

特定のアルゴリズム課題では、検証可能な報酬で学ぶ際の行き詰まりは局所最小値以外の原因からも起こると示した。

何に役立つ?

RLVRで推論課題を学習させるとき、正則化や勾配推定をどう設計するか考える材料になる。

この研究の面白いところ

表形式の学習をスピングラス模型に写像し、理論解析とトランスフォーマーの実験を対応付けた。

どこまで分かった?

地形の厳密な特徴付けは目先の表形式方策と、主に相関のない入力を持つ課題に基づく。一般のRLVR課題すべてが同じ性質を持つとは示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

検証可能な報酬を用いた強化学習(RLVR)は重要だが、それによって新しい推論能力をどこまで学べるかは議論が続いている。本研究は、群や準群の乗算を繰り返すアルゴリズム課題で、RLVRの最適化地形を調べる。そのため、目先の判断だけを行う表形式方策に対するエントロピー正則化付きRLVRを、決定的方策上のエネルギー模型(スピングラス模型)へ写像した。この写像はRLVRで到達できる成果の上界を与え、表形式の設定で最適化地形を厳密に特徴付けられる。 入力間に相関がない広い種類の模型と課題では、理論と実験の両方から、RLVR学習を閉じ込める局所最小値のない穏やかな地形になると示した。実際の難しさは、少なくとも一部は、地形を進む際の拡散的な障壁や勾配推定の誤差などに由来するようだ。これらは解の発見を妨げうる現実的な障害だが、地形そのものが荒いこととは異なる。エントロピー正則化項の選び方で、こうした障害を緩和できる場合が多いことも示した。この理論と整合的に、最後のトークンへの報酬だけを用いて一から訓練したトランスフォーマーが、非可換群の反復乗算に対するアルゴリズム的な思考の連鎖を学習した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Despite the importance of reinforcement learning with verifiable rewards (RLVR), the extent to which it can learn new reasoning capabilities remains debated. Here we study the optimization landscape of RLVR on algorithmic tasks, such as iterated group and quasigroup multiplication. To this end, we map entropy-regularized RLVR over myopic tabular policies onto an energy-based (spin-glass) model over deterministic policies. This mapping upper-bounds what RLVR can achieve, and lets us rigorously characterize the landscape in this tabular setting. We show, both theoretically and experimentally, that for a wide class of models and tasks with uncorrelated inputs, this landscape is benign, containing no local minima that could trap RLVR training. Rather, the practical difficulty of these tasks appears to stem, at least in part, from issues such as diffusive barriers and gradient-estimation error in traversing the landscape. These are genuine obstacles that can prevent a solution from being found, but they are distinct from the landscape itself being rugged. We show that these obstacles can often be mitigated through the choice of entropy regulator. Consistent with this theory, we find that a transformer trained from scratch, using only last-token rewards, successfully learns an algorithmic chain of thought for iterated non-Abelian group multiplications.

arXiv ID: 2609.28625 / 要約の誤りについて