arXiv論文メモ
新着一覧
math.OC · 査読状況未確認

共通の確率的損失を使う二階停留点探索の必要回数

Second-Order Stationarity with Common Random Losses: Matching Tolerance Bounds

Wendao Wu, Haihan Zhang, Chenheng Zhang, Yanyi Li, Chunyuan Zheng, Cong Fang, Haoxuan Li, Zhouchen Lin

この論文をやさしく読む

ひとことで言うと

確率的な非凸最適化で、勾配が小さく曲率も十分良い点を見つけるための必要な問い合わせ回数を解析した研究。

何に役立つ?

二階の最適化アルゴリズムの問い合わせ回数が理論上の限界に近いかを評価する基準になる。

この研究の面白いところ

従来の上界にあった混合項を除き、上界と下界の許容誤差の指数を一致させた。

どこまで分かった?

結論は分散、滑らかさ、次元など要旨にある条件の下でのオラクル計算量であり、実際の実行時間の測定ではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

本研究は、新たなオラクル応答が毎回一つの共通した確率的なスカラー損失の導関数である場合について、確率的な二階停留性を得るための鋭い多項式的な許容誤差の界を示す。勾配とHessian行列がLipschitz連続な母集団目的関数Fに対し、目標は勾配ノルムがε以下で、Hessian行列の最小固有値が−γ以上の点を見つけることである。ε、γは独立した正の許容誤差とする。勾配の分散が有界で、Hessianの誤差がほとんど確実に有界であるとき、新しい勾配またはHessianベクトル積の呼び出しに必要なミニマックス回数は、対数因子を除いてΘ(ε^−3+γ^−5)となる。この特徴付けでは、正のギャップ、滑らかさ、雑音のパラメータを固定し、次元は明示的な多項式の範囲内で増大できる。上界は、従来の新しいHessianベクトル積による保証にあった混合項ε^−2γ^−2を取り除く。ランダムな直線に沿うHessian推定と二進段階の勾配追跡器によって、勾配の変動と、ランダムな符号を持つ曲率の動きを分離する。 下界では、全域で定義された滑らかな確率的損失から両端のコストを実現する。滑らかな分割によって連鎖の長さに比例する負担なしでスカラー雑音を局所化し、厳密な相殺によって応答全体の情報量を制限する。その結果、値、勾配、Hessian行列全体を同時に返し、値の分散が有界な場合でも同じ許容誤差の指数が成り立つ。母集団Hessianが指数ν∈(0,1]のHölder連続性を持つ場合は、対応する次元の条件の下で、新しい勾配・Hessianベクトル積の計算量は対数因子を除いてΘ(ε^−3+γ^−(3+2/ν))となる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We establish tight polynomial tolerance bounds for stochastic second-order stationarity when each fresh oracle response is a derivative of one common random scalar loss. For a population objective $F$ with Lipschitz gradient and Hessian, the target is $\|\nabla F(x)\|\le \epsilon$ and $\lambda_{\min}(\nabla^2 F(x))\ge -\gamma$, with independent tolerances $\epsilon,\gamma>0$. Under bounded gradient variance and almost-surely bounded Hessian error, the minimax number of fresh gradient or Hessian-vector-product calls is $\widetilde{\Theta}\!\left(\epsilon^{-3}+\gamma^{-5}\right)$. The characterization fixes positive gap, smoothness, and noise parameters, suppresses logarithmic factors, and allows dimension to grow within an explicit polynomial envelope. The upper bound removes the mixed term $\epsilon^{-2}\gamma^{-2}$ from the earlier fresh-HVP guarantee. Direct random-line Hessian estimates and a dyadic gradient tracker separate gradient drift from randomly signed curvature motion. The lower bound realizes the endpoint costs through globally defined smooth random losses: a smooth partition localizes scalar noise without a chain-length penalty, while exact cancellation limits the information in the entire response. Consequently, the same tolerance exponents hold even for joint value, gradient, and full-Hessian responses with bounded value variance. For a population Hessian with Hölder exponent $\nu\in(0,1]$, fresh gradient/HVP complexity becomes $\widetilde{\Theta}\!\left(\epsilon^{-3}+\gamma^{-(3+2/\nu)}\right)$ under the corresponding dimension envelope.

arXiv ID: 2609.28238 / 要約の誤りについて