arXiv論文メモ
新着一覧
stat.ML / cs.DS / cs.LG · 査読状況未確認

オンライン逆線形最適化の後悔を次元の平方根に抑える

Tight Regret Bound for Online Inverse Linear Optimization via Multiscale Matrix Weights

Shinsaku Sakaue

この論文をやさしく読む

ひとことで言うと

未知の好みを行動から学ぶ問題で、累積的な選択の損失を次元の平方根程度に抑える理論的手法を示した。

何に役立つ?

オンラインで効用を推定しながら推薦する手法の理論的な性能比較に役立つ。

この研究の面白いところ

時間幅を知らずに働き、既知の下界と定数倍まで一致する後悔上界を得た。

どこまで分かった?

効用と行動が単位球にある条件での理論結果。次元などに関する多項式時間の実装が可能かは未解決である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

固定されているが未知の線形効用を持つオンライン逆線形最適化を研究する。各ラウンドで環境がコンパクトな行動集合を提示し、学習者がそこから行動を推薦すると、環境は同じ集合の中で効用を最大にする行動を返す。効用ベクトルと行動がd次元ユークリッド単位球にあるとき、最適行動に比べた効用の不足を累積した後悔が、どの時間幅についても期待値でO(√d)となる乱択アルゴリズムを与える。時間幅を事前に知る必要はない。Tがd以上のときの既知の下界Ω(√d)から、次元dへの依存は定数倍を除いて最適である。 アルゴリズムは、幾何級数的に並べたスケールごとの多項式特徴空間で、行列の乗法的重みを維持する。線形計画を解いて推薦の分布を選び、利用可能な行動と返された行動を比較してスコア行列を更新する。オラクルの出力とフィードバック行動が有理数なら、線形最適化オラクルに相対的に計算可能な実装でもO(√d)の後悔上界を保つ。次元、時間幅、入力長の多項式時間で同じ達成率を得られるかは未解決である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We study online inverse linear optimization with a fixed unknown linear utility: in each round, an environment presents a compact action set, the learner recommends an action from it, and the environment returns an action that maximizes the utility over the same set. When the utility vector and the actions lie in the $d$-dimensional Euclidean unit ball, we give a randomized algorithm whose regret---the cumulative utility shortfall relative to optimal actions---is $O(\sqrt d)$ in expectation for every time horizon, without knowledge of the horizon. The dependence on $d$ is optimal up to a constant factor by the known $\Omega(\sqrt d)$ lower bound for horizons $T\ge d$. Our algorithm maintains matrix multiplicative weights on polynomial feature spaces at geometrically spaced scales. It selects a recommendation distribution by solving a linear program and updates its score matrices by comparing the available actions with the feedback action. With rational oracle outputs and feedback actions, an implementation computable relative to a linear-optimization oracle preserves the $O(\sqrt d)$ regret bound. Whether the same rate is attainable with running time polynomial in the dimension, horizon, and input length remains open.

arXiv ID: 2609.26978 / 要約の誤りについて