arXiv論文メモ
新着一覧
cs.LG / cs.AI / math.OC · 査読状況未確認

線形スキップ接続がReLU学習の偽の極小を除く条件

Removing spurious minima for planar features by skip connections

Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho

この論文をやさしく読む

ひとことで言うと

ネットワークをただ広くするだけでは消えない学習上の局所極小が、学習可能な線形の迂回経路を加えると消える条件を証明した研究です。

何に役立つ?

スキップ接続が学習のしやすさに寄与する理由を、限定したモデルで数学的に理解する助けになります。

この研究の面白いところ

幅をいくら増やしても偽の極小が残る具体例と、スキップ接続でそれを除ける定理を対比しています。見かけのニューロン数が多くても、異なる特徴方向の数には教師側から上限が付きます。

どこまで分かった?

ガウス母集団損失、浅いバイアスなしReLUモデル、出力重みの符号、教師特徴の平面性などの条件がある理論結果です。任意の深層ネットワークの学習成功を保証するものではありません。経験損失への移行には半径に応じた標本精度が必要です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

損失地形の理解はニューラルネットワークの学習を説明するうえで中心的な課題だが、その構造は単純なモデルでも部分的にしか分かっていない。本研究では教師・生徒設定において、バイアスを持たない浅いReLUネットワークのガウス分布に関する母集団損失を調べる。この設定は、特徴学習や過剰パラメータ化などの本質的な側面を研究するための単純なモデルとなる。出力重みが正で、特徴が平面内にある教師ネットワークについて、学習可能な線形スキップ接続を加えると、生徒ネットワークの幅が教師以上であれば、生徒の出力重みが非負である偽の局所極小がすべて除かれることを示す。対照的に、スキップ接続がない場合には、入力次元が2、隠れニューロンがわずか3個で、出力重みが正の固定された教師ネットワークを構成し、生徒の幅が3以上のどの値でも偽の局所極小が残ることを示す。したがって、学習可能な線形スキップ接続は、任意に過剰パラメータ化しても残る偽の極小を除去できる。 さらに、出力重みが正の生徒ネットワークは常に教師の特徴が張る部分空間を学習することを示す。すなわち、生徒の出力重みが非負である局所極小では、生徒の特徴は教師の特徴が張る空間内にある。2次元のReLUネットワークでは、大幅に過剰パラメータ化された生徒ネットワークであっても、実効的な幅は教師の幅によって制御される。生徒の出力重みが正の任意の臨界点では、生徒の異なる特徴方向の数は、教師ニューロン数の高々2倍である。最後に、この良好な極小構造に関する結果を、任意に指定した半径のパラメータ球上の経験損失の極小へ移し、その際に必要な標本による近似精度が球の半径に依存することを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple models. We study the Gaussian population loss of shallow, bias-free ReLU networks in the teacher--student setting. This provides a simple model for studying essential aspects such as feature learning and overparameterization. For teacher networks with positive output weights and planar features, we show that including a learned linear skip removes all spurious local minima with non-negative student output weights once the student network is at least as wide as the teacher network. In contrast, without the skip, we construct a fixed teacher network with positive output weights and only three hidden neurons in input dimension two whose spurious local minima persist at every student width at least three. Thus, a learned linear skip can remove spurious minima that persist under arbitrary overparameterization. Furthermore, we show that a positive output weight student network always learns the subspace spanned by the teacher features: student features at local minima with non-negative student output weights lie in the span of the teacher features. For ReLU networks in two dimensions, even heavily overparameterized student networks have effective width controlled by the teacher width: every critical point with positive student output weights has at most twice as many distinct student feature directions as teacher neurons. Finally, we transfer the benignity result to empirical minima over parameter balls of any prescribed radius, with the required sampling accuracy depending on that radius.

著者のコメント

43 pages, 4 figures. Under review. Accompanying Lean 4 formalization available at https://github.com/JayPiZimmermann/Removing-spurious-minima-for-planar-features-by-skip-connections

arXiv ID: 2610.01728 / 要約の誤りについて