arXiv論文メモ
新着一覧
stat.ML / cs.LG · 査読状況未確認

ニューラル近似におけるネットワーク幅と重みの大きさ

Optimal Tradeoffs Between Network Size and Parameter Magnitude in Neural Approximation and Minimax Regression

Baicheng Li, Zuowei Shen, Haizhao Yang, Shijun Zhang

この論文をやさしく読む

ひとことで言うと

ニューラルネットワークの幅と重みの大きさをどう組み合わせると、関数近似や回帰で最適な精度が得られるかを理論的に調べた。

何に役立つ?

モデルの大きさとパラメータの許容範囲を選ぶ際の理論的な指針になると考えられる。実データ上の性能を示すものではない。

この研究の面白いところ

幅を増す選択から幅を固定してパラメータの大きさを増す選択まで、同じ最適な統計的収束率を達成する条件を与える。

どこまで分かった?

ヘルダー関数、指定の活性化関数、設計密度と雑音について述べた条件の下での理論結果である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ニューラルネットワークの統計的な精度は、関数を近似する能力と、データに当てはめる関数群の複雑さの両方に依存する。幅を広げるだけでなく、パラメータの大きさも近似と統計の両面で定量化すべき資源である。本研究は、基本的な有界の1リプシッツなDyadic–Triangular活性化関数を用い、深さを固定したときの幅とパラメータの大きさの鋭いトレードオフを示す。[0,1]^d上の単位βヘルダー球(0<β≤1)に対し、幅N≥2d+3、パラメータの絶対値の上限T≥1、0<p<∞のとき、最適なL^p近似誤差のオーダーは [N² log(eNT)]^(−β/d) となる。任意の固定した大域的ヘルダー連続の活性化関数について、対応する下界も成り立ち、そのヘルダー指数は定数には影響するが収束率には影響しない。設計変数の密度が有界で、各誤差が独立・平均0の劣ガウス分布に従う条件では、深さ23のクリップした関数群全体での近似最小二乗法が、N² log(eNT) ≍ M^(d/(2β+d)) のとき、対数損失なしに古典的なヘルダー関数のミニマックスリスク O(M^(−2β/(2β+d))) を達成する。Mは標本数である。これはパラメータの大きさの上限が1の場合からネットワークの大きさを固定する場合まで、統計的に最適な選択の連続体を与える。大きさを固定すると、隠れ層4層・非ゼロパラメータ最大8d+7個でほぼ最適な上限が得られ、6層・最大8d+27個では近似誤差ηに対し log T=O(η^(−d/β)) という最適オーダーに達する。同じ復号方法から、固定サイズのTransformerによる近似も得られる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data. While increasing network size is a natural way to improve approximation, parameter magnitude provides another resource whose role must be quantified in both respects. We establish a sharp width--magnitude tradeoff at fixed depth using one elementary bounded $1$-Lipschitz Dyadic--Triangular Activation. For the unit $\beta$-Hölder ball on $[0,1]^d$ with $0<\beta\leq1$, the optimal $L^p$ approximation error for $0<p<\infty$ is of order $[N^2\log(eNT)]^{-\beta/d}$ when the network width satisfies $N\geq2d+3$ and the parameter magnitudes are bounded by $T\geq1$. Matching lower bounds hold for every fixed globally Hölder activation; its Hölder exponent affects the constants but not the rate. Under bounded design densities and independent centered sub-Gaussian noise, approximate least squares over the full clipped class at depth $23$ attains the classical Hölder minimax risk $\mathcal{O}(M^{-\frac{2\beta}{2\beta+d}})$ without logarithmic loss whenever $N^2\log(eNT)\asymp M^{\frac{d}{2\beta+d}}$, where $M$ is the sample size. This yields a continuum of statistically optimal choices, ranging from unit parameter radius to fixed network size. At fixed size, four hidden layers with at most $8d+7$ nonzero parameters give a near-optimal radius, while six layers with at most $8d+27$ attain the optimal order $\log T=\mathcal{O}(\eta^{-d/\beta})$ at approximation error $\eta$. The same decoding method also yields fixed-size Transformer approximation.

著者のコメント

71 pages

arXiv ID: 2609.25710 / 要約の誤りについて