arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

事前学習モデルで訓練と検証の性能差が広がる仕組み

Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones

Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran

この論文をやさしく読む

ひとことで言うと

学習を続けるうちに、広く役立つ特徴より特定の訓練例にしか役立たない特徴へ更新が偏り、検証性能との差が開くという説明を調べています。

何に役立つ?

検証データを測定指標の入力に使わずに、訓練中の構造変化を観察する方法の検討に役立ちます。

この研究の面白いところ

人工的な特徴階層で条件を操作した実験と、言語・画像モデルでの相関分析を組み合わせています。水準の相関と変化量の関連も分けて報告しています。

どこまで分かった?

実モデルでの相関は、提案した仕組みだけが性能差を生むという因果的な証明ではありません。要旨も、指定した監視指標で観察可能な予測を検証したものと位置付けています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

訓練・検証の分離とは、観測済みの訓練例と有限の未使用検証集合に対する性能の差が、時間とともに変化することである。本研究は、事前学習モデルの適応中にこの差がどう生まれるかについて、動的で構造的な説明を提案する。適合を続けると、更新が必要とされる先が、広く再利用できる特徴の支えから、未使用データへの転用が弱い、より狭い支えへ移り得る。条件付きの局所モデルは、この移行を、勾配配分の不均一性の増加と訓練・検証の分離に結び付ける。固定した訓練プローブにより、測定指標に検証例を入れずにこの構造的な変化を観察できる。未使用データでの性能は、差との関係を評価するために別途用いる。 残差多層パーセプトロン(ResMLP)で実装した人工的な階層では、四段階の共有特徴の相対混合比1:2:3:4を維持しつつ、目標に占める事例固有特徴の割合をp = 0.3、0.5、0.7へ増やすと、各条件5実行での最終平均正答率差が0.185、0.331、0.527へ増加した。訓練例と検証例で別々に測ったマスク入力の損失が、対応する転用の非対称性を示す。 自然言語処理の解析では、RoBERTa、DeBERTa、Qwenを六つのデータセットで10エポック学習した90実行を用いる。訓練プローブで重み付けしたクラス内分散と全体分散の指標は、どちらも全90実行で、未平滑化および平滑化した水準において正答率差と正の相関を持つ。約1エポック間隔で対応付けた未平滑化の変化量も、それぞれ90実行中86実行と87実行で正の関連を保つ。ResNet-18の40エポックの研究では、三つの画像データセットで両指標を検証する。制御されたシミュレーション、自然言語処理、画像の結果を合わせると、この動的で構造的な説明は複数の設定で支持される。実モデルの証拠は、指定した監視指標のもとで、この説明の観察可能な予測を検証している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Train-validation separation is the evolving difference between performance on observed training examples and a finite held-out validation set. We propose a dynamic structural account of how this gap develops during adaptation of pretrained models: continued fitting can shift update demand from broadly reusable support toward narrower support with weaker held-out transfer. A conditional local model links this shift to increasing heterogeneity in gradient allocation and train-validation separation. Fixed training probes make this structural evolution observable without validation examples entering the readouts; held-out performance is used separately to evaluate its relation to the gap. In a constructed hierarchy implemented with a residual multilayer perceptron (ResMLP), increasing the target share of example-private features from $p=.3$ to $.5$ to $.7$, while preserving the relative mixture $1{:}2{:}3{:}4$ among the four shared feature levels, increases the final mean accuracy gap from $.185$ to $.331$ to $.527$ across five runs per condition. Masked-input losses measured separately on training and validation examples expose the corresponding transfer asymmetry. The natural language processing (NLP) analysis uses 10-epoch runs of RoBERTa, DeBERTa, and Qwen on six datasets (90 runs): the training-probe-weighted within-class and overall dispersion readouts each have positive raw and smoothed level correlations with the accuracy gap in all 90 runs. Raw changes paired at approximately one-epoch intervals remain positively associated in 86/90 and 87/90 runs, respectively. A 40-epoch ResNet-18 study tests both readouts on three vision datasets. Together, controlled simulation, NLP, and vision support the dynamic structural account across settings, with real-model evidence testing its observable predictions under the specified monitors.

arXiv ID: 2610.01425 / 要約の誤りについて