到達確率を最適化する省メモリなQ学習
Q-Learning for Reachability in MEC-Free MDPs
この論文をやさしく読む
ひとことで言うと
目標状態へ到達するための行動を、状態間の遷移確率を別途覚えずに学ぶ方法です。特定の構造を持つマルコフ決定過程で、最適な行動方針への収束を保証します。
何に役立つ?
到達目標を仕様として与える意思決定の学習に役立つことが期待されます。状態数に対するメモリの増え方を二次から一次へ抑え、標準ベンチマークでは必要サンプル数も減っています。
この研究の面白いところ
モデルフリーのQ学習に到達可能性の漸近保証を持たせています。対象は限定されていますが、一般のMDPをMEC商で帰着できる構造を利用しています。
どこまで分かった?
保証の対象は非終端の最大エンド成分を含まないMDPです。漸近保証は有限の学習回数での誤差保証とは異なり、要旨には具体的なサンプル数や適用例の詳細はありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
到達可能性の仕様に対する強化学習(RL)は、逐次的な意思決定の基礎となる。先行研究では最適方策への漸近収束が確立されているが、そのために用いられてきたのは、背後にあるマルコフ決定過程(MDP)の遷移確率を明示的に推定しなければならないモデルベースの手法だけである。 私たちは、非終端の最大エンド成分(MEC)を含まないMDPの範囲で、到達可能性について漸近的保証を持つ初のモデルフリー・アルゴリズムQuasarを提案する。この範囲は、標準的なMEC商によってあらゆるMDPを帰着できる基本構成要素である。提案アルゴリズムは古典的なQ学習に従い、時間差分更新を使って、遷移確率を一度も学習することなく最適方策へ収束する。 この学習器は、モデルベース手法が必要とするO(|S|²|A|)のメモリ量を、O(|S||A|)へ削減する。標準化されたQuantitative Verification Benchmark Setでは、従来最先端のモデルベース手法よりも何桁も少ないサンプルで最適方策へ収束する。これらの結果は、到達可能性学習、ひいては仕様に基づく強化学習を実用化するための具体的な一歩となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting learner reduces the memory footprint from the O(|S|^2|A|) that model-based methods require to O(|S||A|). On the standardized Quantitative Verification Benchmark Set, our algorithm converges to the optimal policy with orders of magnitude fewer samples than the previous model-based state-of-the-art. Together these results are a concrete step toward the practical deployment of reachability learning and, with it, of specification-guided RL.
著者のコメント
15 pages, 4 figures
arXiv ID: 2610.01781 / 要約の誤りについて