行動に関する微分を使わず強化学習の方策を改善
FERPO: Forward Entropy-Regularized Policy Optimization
この論文をやさしく読む
ひとことで言うと
行動の良さを予測するモデルの値は使いますが、そのモデルを行動で微分せずに制御方策を学習します。良い行動の候補を複数残すように学習するのも特徴です。
何に役立つ?
連続的な行動を選ぶ強化学習で、価値予測の微分に頼る更新を見直す手掛かりになります。サンプル効率とアクター更新速度の改善がシミュレーション環境で報告されています。
この研究の面白いところ
価値予測の正確さと、その微分の正確さを区別しています。順方向KLを使って複数の有望な行動領域を覆い、同時に元の方策から離れすぎないよう制約する構成です。
どこまで分かった?
評価環境はMuJoCo PlaygroundとManiSkillです。高速化の比較対象はREPPOのアクター更新であり、学習全体が同じ割合で高速化するとは述べられていません。要旨には改善幅や実機での評価はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
連続制御におけるオンライン強化学習の最先端手法のいくつかは、学習したクリティックの行動に関する勾配を使って方策を改善する。しかし、クリティックは通常、リターンを予測するように学習されるため、価値の予測が正確でも行動に関する微分が正確とは限らず、方策更新が不安定な根拠に基づく可能性がある。私たちは、クリティックを行動で微分せず、その価値を用いて方策を改善する、オンポリシーの最大エントロピー強化学習アルゴリズムForward Entropy-Regularized Policy Optimization(FERPO)を提案する。 FERPOは、エントロピーとKullback–Leibler(KL)ダイバージェンスで正則化した方策改善の目的関数から、最適な目標行動分布を導く。続いて、ロールアウト方策から抽出した行動による自己正規化重要度サンプリング(SNIS)で推定した順方向KLの目的関数を最小化し、アクターをこの目標分布に適合させる。KL正則化は、目標分布がロールアウト方策から離れる程度を制限することで、重要度の重みが適切に振る舞うようにする。目標分布の一部のモードを優先し得る逆方向KLの目的関数とは異なり、順方向KLは価値の高い複数のモードを広く覆うことを促し、それによって探索を促進する。 MuJoCo PlaygroundとManiSkillでの実験および要素除去実験では、競争力のある性能とサンプル効率の改善を示した。計算性能のベンチマークでも、Relative Entropy Pathwise Policy Optimization(REPPO)より高速なアクター更新を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).
著者のコメント
Code: https://github.com/Atarilab/FERPO
arXiv ID: 2610.02198 / 要約の誤りについて