arXiv論文メモ
新着一覧
cs.LG / cs.SY / eess.SY · 査読状況未確認

合成データ学習の崩壊を防ぐ人間データ比率を解析

Preventing Model Collapse: A Fisher-Rao Perspective on the Dynamics of Training with Synthetic Data

Matteo Marchi, João Pedro Silvestre, Bahman Gharesifard, Paulo Tabuada

この論文をやさしく読む

ひとことで言うと

AIが作ったデータを繰り返し学習する際に、どの程度の新しい人間データが必要かを、確率分布に適した幾何で理論解析します。

何に役立つ?

合成データと人間由来データを混ぜる際の安定性条件を考える基礎になります。個々のLLMの学習にそのまま使える具体的な比率は要旨には書かれていません。

この研究の面白いところ

通常のユークリッド距離からフィッシャー・ラオ計量へ変えることで、高次元でも意味を失わない評価を目指しています。データ分布の空間そのものの形に着目した方法です。

どこまで分かった?

報告は理論的保証であり、大規模LLMの実学習で特定比率を実証したとの記載はありません。必要比率が以前の示唆と異なると述べていますが、その値や増減の方向は要旨だけでは分かりません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)の大規模化に伴い高品質な人間由来データへの需要が増え続け、その供給を使い尽くしてきたため、現在では合成データを使う学習が日常的に行われている。しかし、合成データによる再帰的な学習はしばしばモデル崩壊を引き起こす。これは、モデルが真の背後のデータ分布を次第に忘れていく、劣化を伴うフィードバックループである。合成データと新たな人間由来データを混ぜて学習することは自然な対策であり、モデル崩壊を防ぐことができる。ただし、学習の安定性を保つために必要な、人間由来データと合成データの比率の正確な最小値は未解決である。 本論文では、モデル崩壊を防ぐために必要な人間由来データの最小比率について、厳密な理論的保証を確立する。先行研究はこの比率の形式的な下限を示していたが、その解析はℝⁿの通常のユークリッド距離に依存し、カテゴリ確率分布の空間に適合していないため、非常に高次元ではその評価が実質的な情報を与えなくなる可能性がある。 そこで本研究では、フィッシャー・ラオ計量の下で過程の動力学を解析し、確率単体の情報幾何的構造を明示的に利用する。次元が増えても安定しており、自明なものにならない、収縮性と不変性の定量的評価を導出する。これにより、モデル崩壊を防ぐために実際に必要なデータ比率が、従来示唆されていたものとは異なることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large Language Models (LLMs) are now routinely trained using synthetic data, since high-quality human data has been exhausted by the ever increasing needs of larger and larger models. However, recursive training on synthetic data frequently induces model collapse, a degenerative feedback loop where models progressively forget the true underlying data distribution. Training on a mixture of synthetic and fresh human data is a logical countermeasure and can prevent model collapse. However, it is an open question as to what is the exact minimum required ratio of human-to-synthetic data to maintain training stability. In this paper, we establish rigorous theoretical guarantees on the minimum rate of human data required to prevent model collapse. Although previous work established a formal lower bound for this ratio, such bound can be vacuous for very high dimensions, as the analysis relies on the usual Euclidean metric in R^n and is not adapted to the space of categorical probability distributions. Instead, in this paper we explicitly leverage the information-geometric structure of the probability simplex by analyzing the dynamics of the process under the Fisher-Rao metric. We derive quantitative contraction and invariance bounds that are stable and do not become trivial as the dimensions increase. Thus, we show that the effective required data ratio to prevent model collapse is different than previously implied.

著者のコメント

8 pages. Extended version of the paper accepted for presentation at the 2026 65th IEEE Conference on Decision and Control (CDC). This version contains the full proofs of the auxiliary lemmas, omitted from the conference version for space

arXiv ID: 2609.18878 / 要約の誤りについて