arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

拡散言語モデルが正しいトークンを書き換えないように学習する

Know When to Hold 'em: Correct-Token Retention in Uniform-State Diffusion Language Models

Mojtaba Nafez, James Henderson

この論文をやさしく読む

ひとことで言うと

何度も文章を書き直す拡散モデルが、間違いだけでなく正しい部分まで変えてしまう問題に対し、正しいトークンを保持する学習損失を加えます。

何に役立つ?

拡散言語モデルの自己訂正で、誤りの修正能力を保ちながら不要な変更を減らすために役立ちます。生成手順を変えず、学習の補助損失で対応しています。

この研究の面白いところ

学習損失の内訳から、未破損位置の誤りがほとんど罰せられない原因を探っています。保持の精度を上げると、修正回数が減るだけでなく、生成の多様性も保たれました。

どこまで分かった?

結果は3モデルと6ベンチマークを中心とする評価です。26.5はパーセントポイントの差です。パープレキシティの改善は文章の事実性や全ての用途での品質を直接保証する指標ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

一様状態拡散モデル(USDM)は、ノイズ除去のどのステップでも任意のトークンを修正でき、自分の誤りを直せる。これはマスク型拡散に対する重要な利点である。しかし自己訂正には、誤ったトークンの修正と正しいトークンの保持の両方が必要であり、現在のUSDMには後者が欠けることを示す。最終段階を貪欲に復号するgreedy-tail復号でも、最先端のUSDMであるDUO、UDLM、一様雑音SEDDは、512位置のうち173〜270位置を各ステップで修正し続ける。この大規模で連携の取れていない編集が、生成サンプルの多様性を崩壊させる。 トークンをランダムに破損させる実験から、この欠陥はモデル自体に由来すると分かる。破損していないトークンはより容易な対象なのに、モデルは未破損と破損したトークンをほぼ同じ精度で再構成する。検証用の負の証拠下界(NELBO)の分解は、学習が保持をほとんど評価していないことを示す。破損位置での誤予測は強く罰せられるが、未破損位置での誤予測はほぼ無罰である。 本研究は、順方向過程で変更されなかったトークンを保持するように学習する、単純だが有効な補助損失、正トークン保持正則化(CTR-Reg)を提案する。サンプラーを変える必要はない。CTR-Regは、6ベンチマークの平均で未破損トークンの精度を26.5パーセントポイント改善する一方、破損トークンの精度はほぼ変えず、各ステップの修正位置はわずか3〜11か所へ収束する。greedy-tailを5ステップ行うだけで、3モデル全ての生成パープレキシティがCTR-Regによって半分未満になり、多様性も維持される。これらの改善は異なるサンプリング予算でも保たれる。結果は、正しいトークンの保持が自己訂正型の拡散言語モデルに欠けていた重要な要素であることを特定し、有効な修正方法を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Uniform-state diffusion models (USDMs) can revise any token at any denoising step, which lets them correct their own mistakes, a key advantage over masked diffusion. Self-correction, however, requires both revising incorrect tokens and retaining correct ones, and we show that current USDMs lack the latter. Even under greedy-tail decoding, state-of-the-art USDMs (DUO, UDLM, and uniform-noise SEDD) keep revising 173--270 of 512 positions at every step, and these large, uncoordinated edits collapse sample diversity. A random-token corruption experiment traces this deficit to the models themselves: they reconstruct clean and corrupted tokens with nearly identical accuracy, even though clean tokens are easier targets. A decomposition of the validation NELBO shows that training barely rewards retention: incorrect predictions are heavily penalized at corrupted positions but almost free at clean ones. We propose Correct-Token Retention Regularization (CTR-Reg), a simple but effective auxiliary loss that trains the model to retain tokens left unperturbed by the forward process and requires no change to the sampler. CTR-Reg improves clean-token accuracy by 26.5 percentage points on average across six benchmarks, while leaving corrupted-token accuracy virtually unchanged, and its per-step revisions converge to only 3--11 positions. With just five greedy-tail steps, generative perplexity more than halves under CTR-Reg for all three models while diversity is preserved, and these gains hold across sampling budgets. Our results identify correct-token retention as a key missing ingredient for self-correcting diffusion language models, and demonstrate an effective fix.

著者のコメント

38 pages, 8 figures

arXiv ID: 2610.01275 / 要約の誤りについて