並列の推測生成を全体情報で修正し採用トークン列を伸ばす
DRelay: Global Draft Context for Prefix-Aware Parallel Speculative Decoding Repair
この論文をやさしく読む
ひとことで言うと
LLMが先にまとめて予想したトークン列を、本体モデルで確認する前に修正し、先頭から連続して使える部分を伸ばす方法です。
何に役立つ?
LLMの生成速度を高める投機的デコードで、並列に出した候補が早い段階の誤選択で無駄になる問題への対策になります。
この研究の面白いところ
候補を増やすだけでなく、ブロック全体の情報とすでに選んだ前半の列を組み合わせて、元の候補を保持するか差し替えるか判断します。先頭から連続して採用される長さを意識して学習損失も重み付けします。
どこまで分かった?
比較結果はH800 GPU上の8ベンチマークとSGLangでの条件に基づきます。要旨には他のハードウェアでの効果や学習コストの具体値は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
並列のドラフト生成は、大規模言語モデル(LLM)の投機的デコードにおけるドラフト生成の負担を減らすが、その効果は採用される接頭列の長さに制約される。正しいトークンが候補集合に含まれていても、早い位置で一度選択を誤るだけで、それ以降の予測を利用できなくなる。そこで、ドラフトブロック全体の情報を使い、対象モデルによる検証の前に、接頭列を考慮して候補選択を選択的に修復するDRelayを提案する。 DRelayは、候補間の相関と選択済みの経路に基づいて判断する。大域的リーダーが、各候補について位置をまたぐ予測情報を抽出する。因果的セレクターは、この大域的な読み取りで得た候補単位の情報を、先行位置で選ばれたトークンと組み合わせ、現在の位置での元の選択が、全体の根拠および選択済み接頭列と整合するかを判断する。そのうえでトークンを保持するか置き換えるかを決め、早い位置の誤りを修復して、採用される接頭列を延ばす。 さらに、候補を支える学習と修復目的を組み合わせて、ドラフトの基盤モデルとセレクターを共同学習する。その際、ブロック内の各位置が連続して採用される接頭列へ寄与し得る度合いに応じて、修復損失を重み付けする。H800 GPU上の多様な8ベンチマークで、DRelayはDFlash、Domino、DSparkを上回り、平均採用長とエンドツーエンドのデコード性能を一貫して改善する。SGLangによるサービングでは、平均エンドツーエンド高速化率を、DFlashに対して14.7〜16.8%、Dominoに対して8.7〜9.3%、DSparkに対して8.1〜9.3%改善する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Parallel drafting reduces the drafting overhead of speculative decoding for large language models (LLMs), but its gains remain limited by the accepted prefix length. Even when the correct token is present in the candidate pool, a single early selection error prevents subsequent predictions from being used. We propose DRelay, which uses global information from the entire draft block to perform prefix-aware selective repair of candidate selections before target-model verification. DRelay bases its decisions on candidate correlations and the selected path: a global reader extracts predictive information across positions for each candidate. While a causal selector combines candidate-level information extracted by the global read with the tokens selected at preceding positions to determine whether the native choice at the current position is consistent with the global evidence and the selected prefix. It then decides whether to retain or replace the token, thereby repairing early errors and extending the accepted prefix. We further jointly train the draft backbone and the selector, combining candidate-support learning with a repair objective, while weighting the repair loss according to each block position's potential contribution to the consecutive accepted prefix. Across eight diverse benchmarks on an H800 GPU, DRelay consistently improves both average acceptance length and end-to-end decoding performance over DFlash, Domino, and DSpark. Under SGLang serving, DRelay improves average end-to-end speedup over DFlash, Domino, and DSpark by 14.7%-16.8%, 8.7%-9.3%, and 8.1%-9.3%, respectively.
arXiv ID: 2610.01439 / 要約の誤りについて