不規則な時点で判断する強化学習の後悔限界
Provably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs
この論文をやさしく読む
ひとことで言うと
判断時点が不規則な連続時間の強化学習で、学習の損失の上下限を示した理論研究。
何に役立つ?
事象に応じて判断する強化学習アルゴリズムの性能保証を考える基礎になる。
この研究の面白いところ
ポアソン過程で決まる判断時点について、二種類のアルゴリズムで同じ次数の保証を得た点。
どこまで分かった?
結果は時間に関する滑らかさなどの仮定の下での理論的保証であり、実環境の性能評価は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
現実の強化学習の問題には連続時間で進み、固定された離散ステップではなく、不規則な事象に応じて判断するものが多い。本研究は、判断時点が一様なポアソン過程に従い、報酬と状態遷移の仕組みが時間とともに滑らかに変わる、エピソード型の連続時間マルコフ決定過程を調べる。各エピソードのジャンプ回数を固定する場合と、時間の上限を固定してポアソン判断時点の回数がランダムになる場合の両方を扱う。時間についてのリプシッツ連続性を仮定し、離散化によって局所的な滑らかさを利用して、UCRLとQ学習をこの設定へ拡張した。モデルに基づく手法とモデルを使わない手法の両方で、対数因子を除きTの3分の2乗の後悔上界を証明した。さらに、同じ次数のミニマックス下界を示し、対数因子を除けばこの達成率が最適であることを示した。これは、ポアソン判断時点を持つリプシッツ連続なエピソード型連続時間マルコフ決定過程について、上下が一致する最初の後悔保証となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 掲載先の記載あり
著者による掲載先の記載:Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:1823-1856, 2026。出版社での独立確認は未実施です。
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Many real-world reinforcement learning (RL) problems evolve in continuous time, where decisions occur at irregular, event-driven intervals rather than at fixed discrete steps. We study episodic continuous-time Markov Decision Processes (MDPs) in which decision epochs are governed by a homogeneous Poisson process and the reward and transition dynamics vary smoothly over time. We consider both a fixed number of jumps per episode and a fixed time budget with a random number of Poisson decision epochs. Under a Lipschitz continuity assumption in time, we exploit local smoothness through discretization and extend both UCRL (Auer and Ortner 2006) and Q-learning (Jin et al. 2018) to this setting, proving $\widetilde{O}(T^{2/3})$ regret bounds for both model-based and model-free algorithms. Finally, we establish matching $\widetilde{\Omega}(T^{2/3})$ minimax lower bounds, showing that the rate is optimal up to logarithmic factors. These results provide the first tight regret guarantees for Lipschitz-smooth continuous-time episodic MDPs with Poisson decision epochs.
arXiv ID: 2609.23127 / 要約の誤りについて