arXiv論文メモ
新着一覧
math.OC · 査読状況未確認

在庫・現金残高の方策最適化で収束率を保証する

Benign Nonconvex Landscape for Policy Optimization: Infinite-Horizon Discounted MDPs with General State and Action Spaces

Xin Chen, Minda Zhao

この論文をやさしく読む

ひとことで言うと

在庫量や現金残高を調整する方策の学習で、非凸な最適化でも悪い停留点にとどまらず、どれくらいの速さで収束するかを理論的に調べています。

何に役立つ?

指定した運用模型で方策勾配法を使う際の、収束を評価する基準になります。従来の方策クラスへの強い閉性条件を弱めることが狙いです。

この研究の面白いところ

Q値関数が完全に凸でなくても、そのずれを停留性の尺度で抑えることで収束条件につなげます。追加の曲率条件によって収束率が変わります。

どこまで分かった?

示される反復計算量と線形収束は、厳密な方策勾配とリプシッツ連続性などの条件付きです。サンプルから推定した勾配で同じ速度が得られるとの保証ではありません。初めてという位置付けは著者らの認識です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

構造を持つ定常方策クラスの下で、一般的な状態・行動空間を持つ無限期間割引マルコフ決定過程(MDP)の最適化の形状を調べる。方策勾配法の大域的収束保証を示す一般的な重み付き方策反復の方法では、どの方策に対しても、重み付き方策改善の結果が同じクラスに含まれるという閉性が必要となる。この性質は、方策クラスが最適方策を含んでいても成り立たない場合がある。 この問題に対処するため、最適でない停留点が存在しないことを保証する、より弱い条件を提案し、有限の集中係数の下で、方策勾配の目的関数に対するPolyak–Łojasiewicz–Kurdyka(PLK)条件を確立する。また、各状態で成立する方策改善の上界から、集中性の仮定なしでもPLK条件を示す。同じ方策クラスとパラメータ領域について共通の基本仮定が成立すれば、本研究の一般的な結果は、従来の枠組みが扱う設定も含む。 さらに、マルコフ変調需要を持つ在庫システムと、確率的な現金残高問題という二つの運用模型で、提案条件を検証する。両模型では、ベルマン方程式から、行動変数に関するQ値関数の近似的な凸性が得られ、その凸性からのずれは一次の停留性の尺度で制御される。これらの評価により指数1のPLK条件が成立し、追加の曲率仮定の下では指数2の条件が成立する。方策勾配のリプシッツ連続性と合わせると、厳密な方策勾配を使う射影勾配降下法について、それぞれO(1/ε)の反復計算量と線形収束が導かれる。著者らの知る限り、方策勾配法で無限期間割引のマルコフ変調需要在庫システムと確率的現金残高問題を解く際の、初めての非漸近的な収束率を与える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We study the optimization landscape for infinite-horizon discounted Markov decision processes (MDPs) with general state and action spaces under structured stationary policy classes. A general weighted policy-iteration approach to establishing global convergence guarantees for policy gradient methods requires closure under weighted policy improvement at every policy, a property that may fail even when the policy class contains an optimal policy. To address this issue, we propose weaker conditions that guarantee the absence of suboptimal stationary points and establish the Polyak--Lojasiewicz--Kurdyka (PLK) condition for the policy gradient objective with a finite concentrability coefficient. We also establish the PLK condition from a policy-improvement bound that holds at every state, without a concentrability assumption. Our general results encompass settings covered by the earlier framework when the common standing assumptions hold for the same policy class and parameter domain. We further verify our proposed conditions for two operations models: inventory systems with Markov-modulated demand and stochastic cash-balance problems. For both models, the Bellman equation yields approximate convexity of the Q-value functions in the action variable, with deviations controlled by the first-order stationarity measure. These estimates establish exponent-one and, under additional curvature assumptions, exponent-two PLK conditions, which, together with Lipschitz continuity of the policy gradient, imply an $\mathcal{O}(1/\epsilon)$ iteration complexity and linear convergence, respectively, for projected gradient descent using exact policy gradients. To the best of our knowledge, we provide the first non-asymptotic convergence rates for solving infinite-horizon discounted inventory systems with Markov-modulated demand and stochastic cash-balance problems using policy gradient methods.

著者のコメント

44 pages, 2 figures

arXiv ID: 2609.21433 / 要約の誤りについて