複数の目的を持つ敵対的バンディットの後悔境界
The Role of Coordinates in Pareto Regret for Adversarial Multi-Objective Bandits
この論文をやさしく読む
ひとことで言うと
複数の報酬を同時に考える逐次選択問題で、最良の後悔の大きさと、それを達成する方法を示す。
何に役立つ?
多目的な逐次意思決定アルゴリズムの理論保証を評価するのに役立つ。
この研究の面白いところ
どの目的が容易か分からなくても適応でき、座標数に依存しない境界を得る点。
どこまで分かった?
選択肢4個以上、試行6回以上、座標2個以上などの条件下での理論結果である。実験的な評価は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
敵対的な多目的バンディットは、報酬が敵対者によって選ばれる多次元ベクトルである選択肢を最適化し、性能をPareto後悔で測る問題を扱う。本研究は損失を1から報酬を引いた値と定義し、ある座標の容易さを、その座標で選択肢の累積損失が最小となる値で測り、その値が小さい座標を容易と呼ぶ。従来の理論は、容易な座標がPareto後悔を減らしうることを示唆するが、実際にはどの座標が容易か分からないことがある。否定的な結果として、著者らは、この情報がないと累積損失が小さくてもPareto後悔の最悪時のオーダーは改善しないことを示す。T回の試行で座標dの最小累積損失をL_dとすると、選択肢が4個以上、試行が6回以上、座標が2個以上の場合、最小最大の期待Pareto後悔はΩ(min{T−L₀, √(K(T−L₀))})であると証明する。この式は、L₀が全座標中の最小値だと分かっている場合にも、L₀が増えるにつれて単調に減る。一方、この結果は、容易な座標だけでなく別の座標からも最適な後悔率を得られる可能性を示す。L₀が既知なら、固定した座標にPoly-INFを適用して、下界と同じオーダーの上界を得る。未知の場合には、報酬を倍増させる方式のPoly-INFを開発し、未知量に適応しながら対応する最小最大の率を達成する。この結果には、余分なlog Tの因子がなく、座標数にも依存しないという含意がある。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Adversarial multi-objective bandits hold the potential to help us optimize choices (arms) whose reward is a multidimensional vector chosen by an adversary and whose performance is measured by Pareto regret. We define loss as one minus reward and measure the easiness of a coordinate by the smallest cumulative loss of the arms on it, and call the coordinate easier when this quantity is smaller. Existing work suggests that in theory an easier coordinate may reduce Pareto regret. However, in practice, one may not know which coordinate is easier. On the negative side, we show that this lack of information eliminates the possibility: a smaller cumulative loss does not improve the worst-case order of Pareto regret. Precisely, let \(L_d\) be the smallest cumulative loss along coordinate $d$ over $T$ rounds. For \(K\ge4\) arms, \(T\ge6\) rounds, and at least 2 coordinates, we prove that the minimax expected Pareto regret is \(\Omega(\min\{T-L_0,\sqrt{K(T-L_0)}\})\). It is monotonically decreasing in \(L_0\), even when \(L_0=\min_d L_d\) itself is known. On the positive side, this result motivates the possibility that other coordinates, not just the easy one, may suffice to attain the optimal rate of Pareto regret. When $L_0$ is known, we apply Poly-INF to a fixed coordinate and obtain an upper bound on Pareto regret that exhibits the same order and thus matches the lower bound. Without such knowledge, we develop a reward-doubling version of Poly-INF that adapts to this unknown quantity while still attaining the matching minimax rate. Another implication is that it has no extra \(\log T\) factor and is independent of the number of coordinates.
arXiv ID: 2609.23092 / 要約の誤りについて