arXiv論文メモ
新着一覧
cs.LG / cs.AI / math.PR / stat.ML · 査読状況未確認

大規模言語モデルの学習と生成を確率過程として整理

The Probabilistic Structure of Large Language Models

Adnan Aboulalaâ

この論文をやさしく読む

ひとことで言うと

LLMの学習と文章生成を確率の言葉で一貫して説明し、拡散モデルとも比較する総説的な論文。

何に役立つ?

LLMの尤度学習、生成、ハルシネーションの議論を共通の数学的枠組みで理解するのに役立つ。

この研究の面白いところ

統計的なもっともらしさと真実を分け、自己回帰生成と逆時間拡散を確率過程として並べて説明する。

どこまで分かった?

要旨は理論的な整理を述べており、新しいモデルの性能評価やハルシネーションを解消する実験は記載していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

本論文は、大規模言語モデル(LLM)を確率論から捉え、文献で別々に扱われがちな道具を一つの自己完結した説明にまとめることを目的とする。LLMを、トークン列の集合上の確率測度として記述し、それを自己回帰的な条件付き分布によって指定する。学習は、確率的勾配法で解く最尤推定問題として定式化し、文章生成は得られた確率過程を順にシミュレートすることとして見る。 また、Kullback–Leiblerダイバージェンスの非対称性が文章生成で果たす役割を、ハルシネーションや、統計的にもっともらしいことと真実であることの違いと関連づけて検討する。同じ見方を補う例として、スコア関数を中心とする拡散モデルも論じる。そこでは生成をトークンの逐次予測ではなく、離散時間と連続時間の双方で、ノイズをデータへ変える逆時間の確率過程のシミュレーションとして捉える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

This paper presents a probabilistic perspective on large language models (LLMs), developed with the aim of bringing together, in a single self-contained account, tools that are usually treated separately across the literature. LLMs are described through probability measures on the set of sequences of tokens, specified via their autoregressive conditional distributions. Training is formulated as a maximum-likelihood estimation problem, addressed by stochastic gradient methods, while text generation is viewed as the sequential simulation of the resulting stochastic process. The role of the asymmetry of the Kullback--Leibler divergence in text generation is examined in relation with characteristic phenomena such as hallucination and the distinction between statistical plausibility and truth. As a complementary illustration of the same viewpoint, we also discuss diffusion models, built around the score function, which cast generation not as sequential token prediction but as the simulation of a reverse-time stochastic process transforming noise into data both in discrete and continuous time.

著者のコメント

Expository article on the probabilistic structure of large language models, 27 pages

arXiv ID: 2609.25134 / 要約の誤りについて