予測期間を知らなくても専門家選択の後悔を抑制
Prediction with Expert Advice: Anytime Regret with Many Experts Matches the Fixed-Time Constant
この論文をやさしく読む
ひとことで言うと
何回予測するかを事前に知らなくても、専門家の助言を使う学習のリグレットを小さく抑える理論研究。
何に役立つ?
オンライン学習で任意の時点に性能保証が必要なアルゴリズム設計の参考になる。
この研究の面白いところ
期間既知の場合に対する係数√2の差が本質的ではないことを示す。
どこまで分かった?
結果はn人の専門家に対する理論的保証。実データ上の性能比較は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
専門家の助言を使う予測は、オンライン学習の基本問題である。期間Tが事前に分かる場合、n人の専門家に対する最小最大の累積リグレットは、漸近的に√(T ln n / 2)となる。Tに合わせて学習率を調整した重みの乗法更新法がこれを達成し、この定数が改善できないことも知られている。 一方、すべての時点tで同時にリグレットの保証が必要な場合、これまで最良の保証は√(t ln n)で、係数が√2だけ悪かった。この差が避けられないかは不明だった。本研究は、避けられることを示す。期間を事前に知る必要がないアルゴリズムを与え、すべてのt≧1について累積リグレットが R_t≦(1+O(√(ln ln n / ln n)))√(t ln n / 2) を同時に満たすことを証明する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Prediction with expert advice is a fundamental problem in online learning. When the time horizon $T$ is known in advance, the minimax cumulative regret over $n$ experts is asymptotically $\sqrt{\frac{T \ln n}{2}}$. This is achieved by the Multiplicative Weights Update algorithm with a learning rate tuned to $T$, and is known to be tight. If instead the regret bound is required to hold simultaneously at every time $t$, the best known guarantee has been $\sqrt{t \ln n}$---a factor of $\sqrt{2}$ worse---and it has remained unknown whether this factor of $\sqrt{2}$ is necessary. We show that it is not. We give an algorithm, requiring no knowledge of the horizon, whose cumulative regret satisfies $R_t \le \bigl(1 + O(\sqrt{\ln \ln n / \ln n})\bigr)\sqrt{t \ln n / 2}$ simultaneously for every $t \ge 1$.
arXiv ID: 2609.27206 / 要約の誤りについて