arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

学習後に遅れて汎化する現象を重み減衰と特徴学習で説明

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands

この論文をやさしく読む

ひとことで言うと

訓練問題は解けるのに新しい問題は解けない期間が続き、後になって急に汎化する現象を、学習中の特徴の変化から説明する研究です。

何に役立つ?

学習率と重み減衰を変えたとき、汎化がいつ起こり、どの条件で起こらないかを理解するために役立ちます。

この研究の面白いところ

訓練正解率が頭打ちでも内部の課題に関係する構造は育ち続けると捉え、汎化までの時間を学習率と減衰の積で説明しています。

どこまで分かった?

理論は二乗損失を使う斉次ネットワークを対象とし、実験課題は剰余加算です。Transformerでは類似の挙動を観察していますが、理論の厳密な斉次性の仮定は満たしません。格子の大きさは学習条件数で、ネットワークの層数ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

グロッキングでは、訓練データへの早い段階での適合と、そのはるか後に起こる汎化性能の改善が時間的に分かれる。この遅延の間、学習は、ニューラル・タンジェント・カーネル(NTK)が固定された領域から、課題に関連するカーネルの固有方向が変化し続ける領域へ移り得る。本研究では、この怠惰な学習から豊かな学習への移行が、どのように汎化の遅延を生むかについて定量的な理論を与える。 二乗損失とL₂重み減衰で学習する斉次ネットワークでは、訓練データを記憶した後にも有限の残差が残り、NTKの固有値が小さい目標成分ほど残差の割合が大きくなることを示す。これらの残差はNTK自身のダイナミクスにフィードバックする。その結果のダイナミクスを課題に関連するスペクトル方向へ射影すると、残差に駆動されるカーネルの成長と重み減衰が競合する縮約系が得られる。この系は、グロッキングの時間尺度が学習率と重み減衰の積に制御されること、ある臨界減衰値に近づくと特徴学習が対数的に遅くなること、臨界値を超えると課題に整合したNTK構造が汎化を支えられなくなること、さらに強い減衰は適合そのものを妨げ得ることを予測する。 これらの予測を剰余加算で検証する。斉次な多層パーセプトロン(MLP)では、訓練正解率が飽和した後も、課題に整合したフーリエ構造がNTK内に現れ続ける。学習率と重み減衰を変えた84×90の格子状の条件で学習したネットワークから、予測された相の形状と、汎化までの時間が学習率と重み減衰の積の逆数に比例する関係が再現される。1ブロックのTransformerでも、42×45の格子状の条件で同様の巨視的な相構造が見られ、厳密な斉次性を満たさないにもかかわらず、同じ遷移時間のスケーリングが得られる。これらの結果は、適合後の特徴学習を、汎化の始まりと、学習率・重み減衰平面上の相構造の両方に結び付ける機構的な導出を与える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

In grokking an early fit to the training data separates from a much later improvement in generalization. During this delay, training can move from a fixed neural tangent kernel (NTK) regime to one in which task-relevant kernel eigendirections continue to evolve. We provide a quantitative theory for how this transition from lazy to rich learning can produce delayed generalization. For homogeneous networks trained with squared loss and $L_2$ weight decay, we show that a finite residual remains after memorization, with larger residual fractions in target components associated with smaller NTK eigenvalues. These residuals feed back into the dynamics of the NTK itself, and projecting the resulting dynamics onto task-relevant spectral directions yields a reduced system in which residual-driven kernel growth competes with weight decay. This system predicts that the grokking timescale is controlled by the product of learning rate and weight decay, that feature learning slows logarithmically near a critical decay above which task-aligned NTK structure can no longer support generalization, and that stronger decay can prevent fitting altogether. We test these predictions in modular addition. In a homogeneous MLP, task-aligned Fourier structure continues to emerge in the NTK after training accuracy has saturated, and an 84$\times$90-grid of trained networks across varying learning rate and weight decay recovers the predicted phase geometry and inverse-product scaling of the generalization time with learning rate and weight decay. A one-block Transformer shows similar macroscopic phase structure in a 42$\times$45-grid, as well as the same transition-time scaling despite violating exact homogeneity. Together, these results provide a mechanistic derivation connecting post-fit feature learning to both the onset of generalization and its phase structure in the learning rate and weight decay plane.

arXiv ID: 2609.26679 / 要約の誤りについて