arXiv論文メモ
新着一覧
math.ST / cs.LG / stat.ML / stat.TH · 査読状況未確認

離散拡散モデルの必要データ数を局所依存から評価する

Sample complexity bounds for categorical Markov random fields via Discrete Diffusions

Shivam Kumar and Nabarun Deb

この論文をやさしく読む

ひとことで言うと

局所的な依存を持つ離散データを生成する拡散モデルについて、必要な学習データ数から生成誤差までを理論で結び付けます。

何に役立つ?

離散データの生成で、データ量・語彙数・相互作用の複雑さが精度にどう影響するかを見積もる基礎になります。学習後に計算予算に応じて生成手順を調整できます。

この研究の面白いところ

時間依存と対象分布依存を分ける分解を使い、複数のノイズ水準で重みを共有します。学習誤差を既知と仮定せず、有限のデータから解析する点も特徴です。

どこまで分かった?

理論は低次MRFで表される局所依存と一様なノイズ付加が前提です。数値実験はPotts、Ising、木構造モデルであり、例として挙げたすべての実応用での性能検証ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

統計学、経済学、物理学の多くの応用では、局所的な依存構造を持つ高次元のカテゴリカル分布からのサンプリングが必要となる。例として、有限の記憶を持つ言語モデル、統計物理学のIsing系やPotts系、タンパク質の折り畳みなどがある。現代の機械学習では、離散拡散がこうしたデータをサンプリングする柔軟な方法として登場し、高い実証性能を示している。 これを動機として、低次のMarkov確率場(MRF)でモデル化した局所依存の下で、一様なノイズ付加を用いる離散拡散について、学習からサンプリングまでを通した標本複雑度の上界を持つ学習法を開発する。主な技術的着想は、離散スコアの新しい「ピン留め分解」である。連続拡散とは異なり、スコアを、時間への依存と対象分布への依存が乗法的に分離される成分へ分解できることを示す。 この分解に基づいて重み共有型のニューラルスコア学習器を提案し、τ-leapingと組み合わせて、一貫したサンプリング手順を得る。既存のサンプリング解析で一般的なようにスコア学習の誤差を外部から与えられる未知の量として扱うのではなく、有限データから生じるスコア学習誤差を調べ、語彙数、MRFの相互作用次数、標本数への依存を明示した最適なサンプリング保証を導く。 さらに、一様ノイズの各水準を通じて単一のスコアネットワークを学習し、サンプリングの離散化は推論時に選べるようにする。これにより、同じ学習済みモデルで、推論の予算に応じて精度と計算費用を調整できる。Potts、Ising、木構造のモデルでの数値実験では、長い系列のサンプリングにおいて、重み共有型スコアネットワークが全結合型を上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Many applications in statistics, economics, and physics require sampling from high-dimensional categorical distributions with local dependence structures. Examples include finite memory language models, Ising and Potts systems in statistical physics and protein folding, etc. In modern machine learning, discrete diffusions have emerged as a flexible approach for sampling such data, with strong empirical performance. Motivated by this, we develop learning methods with end-to-end sample complexity bounds for discrete diffusion with uniform noising under local dependence, which we model through low order Markov random fields (MRFs). Our main technical insight is a new \emph{pinning decomposition} of the discrete score. It shows that unlike in continuous diffusions, the score decomposes into components where the dependence on time separates multiplicatively from the dependence on the target. Building on this decomposition, we propose a \emph{weight-sharing neural score learner} and combine it with $\tau$-leaping to obtain an end-to-end sampling procedure. Rather than treating score-learning error as a black-box input, as is common in existing sampling analyses, we study the score learning error from finite data and derive optimal sampling guarantees with explicit dependence on the vocabulary size, the interaction order of the MRF, and the sample size. Moreover, our strategy trains a single score network across uniform noise levels while leaving the sampling discretization to be chosen at inference-time. This allows the same trained model to trade accuracy for computational cost as inference-time budgets vary. Numerical experiments on Potts, Ising, and tree-structured models show that weight-sharing score networks outperform fully connected ones for sampling long sequences.

著者のコメント

83 Pages, 3 Figures, 4 Tables

arXiv ID: 2610.02128 / 要約の誤りについて