arXiv論文メモ
新着一覧
stat.ML / cs.LG / math.ST / stat.TH · 査読状況未確認

高次元のバンディット問題で選択法ごとの損失を厳密比較

Exact Bayes Regret and Asymptotic Optimality in High-Dimensional Gaussian Bandits

Prakhar Singhvi, Yi Zou, and Abhishek Bhattacharjee (Abstract Math Institute)

この論文をやさしく読む

ひとことで言うと

試しながら最良の選択肢を探す高次元の問題で、複数の選び方による損失を数学的に比較した研究。

何に役立つ?

探索と利用の方策を選ぶ際に、情報収集量と意思決定の質を分けて評価する理論的な基準になる。

この研究の面白いところ

この設定では事後平均による貪欲法が極限で最適となり、トンプソン・サンプリングの瞬間的な損失は長時間で約2倍に近づく。

どこまで分かった?

結論はガウス分布のパラメータ・候補行動・雑音と、時間が次元に比例する設定での極限結果である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

時間の長さが次元に比例する場合について、パラメータが等方的なガウス分布に従い、候補となる行動が独立したガウス分布から生じ、報酬雑音もガウス分布に従うベイズ線形バンディットを研究する。正規化した事後不確かさには、すべての因果的な方策に対して一様に成り立つ明示的な極限がある。次に、ガウス分布の事後確率に関する恒等式から、適応的な再帰計算に追加の閉じた形を仮定せずに、パラメータの重なりの極限を決める。これにより、トンプソン・サンプリング、事後平均による貪欲な選択、事後サンプリングの共分散を拡大縮小する方策群について、厳密なリグレット曲線を得る。正規化した実現累積リグレットは、次元に比例する時間のコンパクトな区間上で一様に、L¹の意味で収束する。 方策によらない下界から、最適なベイズ・リグレットの極限を特定し、事後平均による貪欲な選択がそれを達成すると証明する。トンプソン・サンプリングの主要項のリグレットはこれより厳密に大きい。瞬間的なリグレットの貪欲法に対する比は1と2の間で、次元に比例する長い時間の極限では2に近づく。累積リグレットの閉形式の曲線は、雑音がゼロに近づく極限では別の比較関係も明らかにする。最後に、瞬間的なリグレットはその平均値ではなく、退化していないガウス型の意思決定損失分布へ収束する。この解析は、方策が得る情報の量と、その情報を使った判断の質を分けて考える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We study Bayesian linear bandits with an isotropic Gaussian parameter, independent Gaussian candidate arms, and Gaussian reward noise when the horizon is proportional to the dimension. The normalized posterior uncertainty has an explicit limit that is uniform over all causal policies. Gaussian posterior identities then determine the limiting parameter overlaps without an assumed closure of the adaptive recursion. These results yield exact regret curves for Thompson sampling, posterior-mean greedy selection, and a family of policies that scale the posterior sampling covariance. The normalized realized cumulative regret converges in L1, uniformly on compact proportional-time intervals. A policy-uniform lower bound identifies the limiting optimal Bayes regret and proves that posterior-mean greedy selection attains it. Thompson sampling incurs a strictly larger leading regret; its instantaneous regret ratio relative to greedy selection lies between one and two and approaches two at long proportional horizons. Closed-form cumulative curves also identify a different comparison in the vanishing-noise limit. Finally, the instantaneous regret converges to a nondegenerate Gaussian decision-loss distribution, rather than to its mean. The analysis separates the amount of information acquired by a bandit policy from the quality of the decisions made using that information.

著者のコメント

17 Pages

arXiv ID: 2609.28718 / 要約の誤りについて