arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

リスクと価値推定の一致を使う頑健な敵対的強化学習

Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization

Jiaxi Wu, Tiantian Zhang, Yuxing Wang, Yongzhe Chang, Xueqian Wang

この論文をやさしく読む

ひとことで言うと

敵対的な摂動の強さを状態に応じて調整し、二つの価値推定器を安定させる。

何に役立つ?

考えられる用途は、不確実性のある連続制御での強化学習の頑健化である。

この研究の面白いところ

強すぎる敵対者が学習を妨げる点と、価値推定器の不一致を同時に扱う。

どこまで分かった?

連続制御ベンチマークでの改善を報告する。実環境の制御結果は要旨にない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

強化学習は連続的な意思決定で高い性能を示すが、動的な不確実性や分布の変化には脆い。頑健な敵対的強化学習(RARL)は最悪条件の摂動で頑健性を高めるものの、最適化が不安定になり、価値推定が悪化しやすい。敵対者が強すぎると、エージェントを学習に役立たない失敗状態へ追い込み、摂動が二つの価値評価器の不一致を広げて、偏った価値目標を生む。提案するRACERは、敵対的強化学習をリスク感度の観点から見直す統一的な枠組みである。まず、状態に応じて摂動の強さを調節する敵対的目標を導入し、有害な妨害を抑えながら有益な探索を保つ。次に、Q値推定器間の不一致を減らして学習を安定させるため、評価器の一致を促す正則化を提案する。難しい連続制御ベンチマークでの包括的な実験では、強力な頑健強化学習の比較手法より、性能、頑健性、学習安定性が一貫して改善した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Reinforcement learning (RL) achieves strong performance in sequential decision-making but remains brittle under dynamic uncertainty and distributional shifts. Robust Adversarial Reinforcement Learning (RARL) improves robustness via worst-case perturbations, but existing approaches frequently suffer from unstable optimization and degraded value estimation. In particular, overly aggressive adversaries can drive the agent toward uninformative failure states, while adversarial perturbations amplify disagreement between double critics and introduce biased value targets. We propose a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective. First, we introduce a state-dependent adversarial objective that adaptively regulates perturbation strength, suppressing harmful disturbances while preserving informative exploration. Second, we propose critic consistency regularization to reduce disagreement between Q-value estimators and stabilize learning. Comprehensive experiments on challenging continuous control benchmarks demonstrate that RACER consistently improves performance, robustness, and training stability over strong robust RL baselines.

arXiv ID: 2609.27667 / 要約の誤りについて