初期状態で最適報酬が変わる頑健Markov決定過程の理論
Vector Bellman Theory for Multichain Robust Average-Reward Markov Decision Processes
この論文をやさしく読む
ひとことで言うと
初期状態によって長期報酬が異なりうる頑健な意思決定問題を、ベクトル型の方程式で扱う理論です。
何に役立つ?
遷移に不確かさがある長期計画で、最適方策や収束を検証するための理論に役立ちます。
この研究の面白いところ
すべての初期状態を同時に扱う鞍点戦略と、有限の計算後に最適となる計画法を結び付けています。
どこまで分かった?
結論は要旨にある有限モデル、曖昧さの形、Bellman解の存在などの条件の下で述べられています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
頑健な平均報酬Markov決定過程は、不確かさの下で長期的な性能を最適化する基本的な枠組みであり、最適な長期報酬が初期状態に依存する場合がある。この依存性を扱うには、再帰的に訪れる状態群の報酬と遷移の不確かさの双方を考慮した、ベクトル型Bellman理論が必要となる。この研究は、行動後に定まる状態・行動ごとの矩形型でコンパクトな曖昧さを持つ有限モデルについて、その理論を構築する。利得を先に、バイアスを次に最適化する原理から、結合したベクトル利得・バイアス系を得る。その有限解はすべて、最適な頑健利得を特定し、履歴に依存する相手に対しても、あらゆる初期状態から同時に使える定常的な鞍点戦略を与える。 さらに、定常的な利得条件と、標準的な一時的補正への一様な上界を使い、解が存在する条件を特徴付ける。異なる再帰クラスが異なる利得を持てる十分条件も示す。この証明用の条件から、頑健Bellman作用素の軌道が漸近的にアフィンな形となることも得られる。それに基づき、頑健な近似シフト付きHalpern計画アルゴリズムを設計する。有限なBellman解が存在すれば、利得推定値とBellman変位は最適利得ベクトルへ収束し、取り出される貪欲な制御方策は、問題例に依存する有限の計算予算の後に平均報酬で最適になる。この結果は、有限Bellman条件と、状態依存の頑健平均報酬のための割引を用いない計画法を結び付ける。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Robust average-reward Markov decision processes provide a fundamental framework for long-term performance optimization under uncertainty, and can have optimal long-run rewards that depend on the initial state. This state dependence requires a vector Bellman theory that accounts for both recurrent-class rewards and transition uncertainty. We develop such a theory for finite models with compact, post-action $(s,a)$-rectangular ambiguity. A gain-first, bias-second optimization principle yields a coupled vector gain-bias system, and every finite solution identifies the optimal robust gain and supplies stationary saddle strategies against history-dependent opponents, simultaneously from all initial states. We further characterize solvability through stationary gain conditions and a uniform bound on canonical transient corrections, and give sufficient conditions that permit distinct recurrent-class gains. The certificates also yield asymptotically affine trajectories of the robust Bellman operator, based on which we design a robust approximately shifted Halpern planning algorithm. Under finite Bellman solvability, the gain estimates and Bellman displacements converge to the optimal gain vector, and every extracted greedy controller is average-optimal after a finite, instance-dependent budget. These results thus connect finite Bellman certificates to undiscounted planning for state-dependent robust average rewards, providing theoretical understandings.
著者のコメント
preprint, work in progress
arXiv ID: 2609.28792 / 要約の誤りについて