arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

拡散型言語モデルの誤答固定を途中状態から修正する

LOCKR: A Hidden-State Trajectory-Guided Planner for Detecting and Repairing Stable-but-Wrong Lock-In in Diffusion Language Models

Guoshenghui Zhao, Tan Yu, Weijie Zhao

この論文をやさしく読む

ひとことで言うと

拡散型言語モデルが途中で誤答に固定されたかを隠れ状態の変化から見つけ、追加計算で修正する方法を評価した。

何に役立つ?

考えられる用途は、推論が誤答に固定されている場合にだけ追加の計算を配分すること。示された改善は数学的推論の評価設定での結果である。

この研究の面白いところ

最終回答の確信度だけでなく生成途中の軌跡を使い、修正が必要かどうかと、どの修正分岐を採るかの両方を判断する。

どこまで分かった?

評価は2モデル、3つの数学的推論ベンチマーク、計5設定に限られる。ほかの課題やモデルで同じ改善幅になるかは要旨からは分からない。

v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

拡散型言語モデルは反復的なノイズ除去によって文章を生成するため、最終回答に至るまでの中間的な軌跡を観察できる。本研究は、まだ多くのノイズ除去が残っている段階で回答が誤った値のまま安定する「安定しているが誤り」という固定化を、繰り返し現れる推論上の失敗として特定する。確信度、エントロピー、回答候補間の差、回答の安定性といった表面上の復号信号だけでは、正しい固定化と誤った固定化を信頼して区別できない。 そこで、必要な場合に限って推論を修正することを、推論時の軽量な計画問題として定式化し、隠れ状態の軌跡を使うプランナーLOCKRを提案する。LOCKRは追加計算を割り当てる時点を決め、狙いを定めた修正分岐を構造的に展開し、軌跡を考慮した検証によって最も有望な続きを選ぶ。2つの拡散型言語モデルと3つの数学的推論ベンチマークで、隠れ状態の軌跡は、誤答の固定化の検出でも修正先の選択でも、表面信号や隠れ状態の単一時点の情報を一貫して上回った。自然な評価分布では、評価した5設定すべてで正答率が絶対値で2.21~5.37パーセントポイント上がり、修正率は22~41%だった。これらの結果は、拡散過程の隠れ状態の軌跡が、必要な場合だけ推論を修正するための実用的な信号となることを示す。

v2の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-24 · v2
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Diffusion language models generate text through iterative denoising, exposing intermediate trajectories before final answers are produced. We identify a recurring reasoning failure, stable-but-wrong lock-in, where an answer stabilizes early around an incorrect value while substantial denoising remains. Surface-level decoding signals such as confidence, entropy, margin, and answer stability are insufficient to reliably distinguish correct from erroneous lock-in. We formulate selective reasoning repair as a lightweight test-time planning problem and propose LOCKR, a hidden-state trajectory-guided planner that decides when to allocate additional computation, expands a structured set of targeted repair branches, and selects the most promising continuation using trajectory-aware verification. Across two diffusion language models and three mathematical reasoning benchmarks, hidden-state trajectories consistently outperform surface signals and single hidden snapshots for both wrong-lock-in detection and repair selection. On natural evaluation distributions, LOCKR yields absolute accuracy gains of 2.21--5.37 percentage points across all five evaluated settings, with repair rates ranging from 22% to 41%. These results establish hidden diffusion trajectories as actionable signals for selective test-time reasoning repair.

著者のコメント

9 pages, 6 figures, appendix included

arXiv ID: 2609.27220 / 要約の誤りについて