arXiv論文メモ
新着一覧
cs.LG / cs.AI · 査読状況未確認

内部状態を使って拡散言語モデルの生成を導く

How to Guide Your Language Flow

Rohit Dilip, Tianrong Chen, Yuyang Wang, David Van Valen, Joshua Susskind, Miguel Angel Bautista

この論文をやさしく読む

ひとことで言うと

拡散型の言語モデルがすでに持っている内部状態から、生成を良い方向へ導く信号を取り出す方法です。

何に役立つ?

考えられる用途は、推論時の追加の順伝播を避けながら生成品質を改善することです。無条件生成と、17億パラメータモデルの多肢選択質問応答で改善が報告されています。

この研究の面白いところ

誘導用に弱いモデルと強いモデルを用意する考え方を、固定された内部状態へのプローブで実現しようとしています。どの訓練段階のモデルが誘導に適するかも調べています。

どこまで分かった?

要旨には具体的な性能差や評価セット名がありません。また「強いモデルが弱いチェックポイント」という原文の記述は関係が分かりにくく、強弱の対応の詳細は要旨だけでは確認できません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

フローマッチングモデルを誘導する新たな方法を導入する。プローブ誘導と呼ぶ本手法は、既存の拡散モデルの固定された内部状態を使って誘導信号を構成する。自動誘導と似た原理で働くが、推論時の追加の順伝播を不要にし、弱いモデルと強いモデルが似た動力学を共有することを確保するための信頼できる方法を提供する。 連続拡散言語モデルにこの手法を適用してベンチマーク評価を行い、プローブ誘導が無条件生成で新たな最高性能を達成した。17億パラメータの拡散言語モデルへ適用すると、多肢選択式質問応答のベンチマークでも一貫して改善した。 また、プローブを用い、強いモデルが弱いチェックポイントである従来の自動誘導の設定を調べたところ、弱いモデルは訓練中の低エントロピーの領域から得られたものでなければならないことが分かった。これらの知見は、拡散言語モデルを改善する実用的な方法を与えるとともに、現時点では十分に理解されていない自動誘導の実際の仕組みに光を当てる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We introduce a new method to guide flow matching models. Our approach, which we call probe guidance, uses the frozen internal states of an existing diffusion model to construct a guidance signal. This works using a similar principle as autoguidance, but eliminates the need for an additional forward pass at inference time and provides a reliable path to ensure that the weak and strong model share similar dynamics. We apply and benchmark this method on continuous diffusion language models, where probe guidance sets a new state-of-the-art performance on unconditional generation. When applied to a 1.7B diffusion language model, probe guidance consistently improves on multiple choice question answering benchmarks. Using our probes, we study the traditional autoguidance setting where the strong model is a weak checkpoint, and find that the weak model must come from a low-entropy region of training. These findings both provide a practical way to improve diffusion language models and shed light on the actual mechanism behind autoguidance, which is currently poorly understood.

arXiv ID: 2609.19356 / 要約の誤りについて