回帰の近似誤差と観測ノイズを明示的な上界で評価
User-friendly approximation theory for statistical regression
この論文をやさしく読む
ひとことで言うと
回帰で生じる「モデルが関数を表し切れない誤差」と「観測ノイズの誤差」を、同じ枠組みで見積もる理論です。
何に役立つ?
モデルの種類、標本の取り方、ノイズの仮定から、回帰誤差の具体的な保証を組み立てるのに役立ちます。定数を明示することで次数だけの評価より計算に使いやすくしています。
この研究の面白いところ
特定の標本格子を使うFourier・Chebyshev回帰では、誤差増幅を表すLebesgue定数の上界をO(√m)からO(log m)へ改善しています。
どこまで分かった?
O(log m)は記載された回帰と標本格子におけるLebesgue定数の上界です。すべての回帰誤差がこの次数になるという主張ではなく、ノイズについても所定の裾・モーメント条件があります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ノイズを含む関数値を用いた回帰の誤差は、モデル空間の選択による近似誤差と、観測ノイズによる確率的誤差に自然に分解される。統計理論は主に確率的誤差を扱う一方、近似理論は近似誤差に焦点を当て、しばしば観測にノイズがないことを仮定する。本研究は、近似理論と統計学の道具を組み合わせ、一般化最小二乗回帰で両方の誤差に上界を与える枠組みを構築する。結果は、重み付き上限ノルムとL²(μ)ノルム、実数値関数と複素数値関数、さらに裾やモーメントに関するさまざまな仮定の下での相関ノイズを扱う。重み付きLebesgue関数とLebesgue定数が最良近似誤差の増幅を制御し、集中不等式が確率的誤差の上界を与える。これらの結果を、定数を明示した再利用可能な上界として整理し、可能な場合にはその鋭さも確立する。 一般的なランダム化回帰の戦略について、Lebesgue関数と定数が、対応するL²(μ)射影の値の周辺に集中することを証明する。標準的な標本格子上でのFourier回帰とChebyshev回帰では、モデル次元をmとして、Lebesgue定数に対する明示的かつ漸近的にタイトなO(log m)上界を得る。これは回帰の設定で従来得られていたO(√m)上界を改善し、補間について知られている上界と一致する。近似誤差の上界の一覧と具体的な計算例により、これらの結果を回帰に対する明示的な保証へ変換する方法を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The error of regression using noisy function values naturally decomposes into an approximation error from the choice of model space and stochastic error from the noise on the observations. Statistical theory is mostly concerned with the stochastic error while approximation theory focuses on the approximation error, often subject to the assumption of noiseless observations. We combine the tools from approximation theory and statistics into a framework for bounding both errors in generalized least squares regression. The results cover weighted sup-norms and $L^2(\mu)$-norms, real- and complex-valued functions, and correlated noise under a range of tail and moment assumptions. Weighted Lebesgue functions and constants control the amplification of the best approximation error, while concentration inequalities provide bounds on the stochastic error. We organize these results into reusable bounds with explicit constants and establish sharpness where possible. For the common strategy of randomized regression we prove concentration of the Lebesgue function and constant around those of the corresponding $L^2(\mu)$-projection. For Fourier and Chebyshev regression on standard sampling grids, we obtain explicit, asymptotically tight $\mathcal O(\log m)$ bounds on the Lebesgue constants where $m$ is the model dimension. This improves upon the previous $\mathcal O(\sqrt{m})$ bound obtained for the regression setting and matches the bounds known for interpolation. A catalogue of approximation bounds and worked examples illustrates how to turn these results into explicit regression guarantees.
arXiv ID: 2610.01342 / 要約の誤りについて