データ繰り返しとモデルスパース性の関係
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
短い要約(全文訳を準備中)
データ繰り返しによるモデルの過学習に焦点を当て、スパースなMixture-of-Expertsモデルの性能低下を調査。データ繰り返し率とモデル構造の影響を分析し、正則化手法の効果を検証。モデルの専門化と過学習の関連性を明らかに。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-10(UTC)
- 最新改訂
- 2026-09-10 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-10 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8x with minimal degradation, MoEs instead begin to suffer at 4x, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to underperform dense models after 32x. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
arXiv ID: 2609.11917 / 要約の誤りについて