繰り返す預言者不等式で後悔を抑える学習法
Optimal No-Regret Learning for Repeated Prophet Inequality
この論文をやさしく読む
ひとことで言うと
順番に現れる未知の値から一つを選ぶ課題を繰り返し、最適な停止方法を学ぶ研究。
何に役立つ?
観測できるのが選んだ位置まででも、試行回数に応じた後悔を小さくするアルゴリズム設計に役立つ。
この研究の面白いところ
観測の入れ子構造を利用し、箱の数への多項式依存なしで、下界にほぼ一致するÕ(√T)の後悔を達成した。
どこまで分かった?
値が独立に引かれ、値域が[0,1]で、箱の順番が固定された理論設定である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
接頭部分だけが観測できる繰り返し預言者不等式を研究する。T回の各ラウンドで、学習者は固定の順番に並ぶn個の箱から、未知で値域が[0,1]の分布に従って独立に新しく引かれた値に出会う。一つを取り消し不能な形で受け入れなければならず、停止した箱までの接頭部分しか観測できない。後悔は、分布を知っている最適停止方策に対して測る。 期待後悔が対数因子を除いて下界に一致する、Õ(√T)の効率的なアルゴリズムを与える。アルゴリズムは、経験的な後ろ向き帰納法と箱ごとの到達ボーナスを組み合わせ、ほぼ最適な方策を直接探索する。さらに、観測される接頭部分の入れ子構造を利用する相対的な減少の集約規則で探索を保ち、箱の数nへの多項式的な依存をなくす。これにより、Liuら(2025年)が提起した未解決問題に答える。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We study repeated prophet inequalities under prefix feedback. In each of $T$ rounds, a learner encounters fresh values drawn independently from $n$ boxes with unknown $[0,1]$-supported distributions in a fixed order and must irrevocably accept one, observing only the prefix up to its stopping box. Regret is measured against the optimal stopping policy that knows the distributions. We give an efficient algorithm achieving $\widetilde O(\sqrt{T})$ expected regret, matching the lower bound up to logarithmic factors. Our algorithm explores directly through near-optimal policies, combining empirical backward induction with box-specific reach bonuses. A relative-drop aggregation rule then exploits the nesting structure of observed prefixes to preserve exploration, thereby removing the polynomial dependence on the box number $n$. This resolves an open question posed by Liu et al. (2025).
arXiv ID: 2609.23265 / 要約の誤りについて