arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

反復型Transformerの途中予測を使って復号を改善

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

この論文をやさしく読む

ひとことで言うと

同じ計算ブロックを何度も通すモデルで、途中の予測と最後の予測の差を利用して次のトークンを選びます。これまで捨てていた中間状態を使う方法です。

何に役立つ?

反復型の言語モデルの精度と推論計算量を改善するために役立ちます。評価では反復回数を半分にしても元の全深さの性能を維持または上回る結果が示されています。

この研究の面白いところ

弱い補助モデルを別に用意する必要がなく、同一モデルの早い段階を対比相手にします。隠れ状態を使う方式では、出力計算自体を追加しません。

どこまで分かった?

対象は4系列のループ型Transformerです。22.5〜48.2%は順伝播FLOPsの削減で、実測の待ち時間短縮率とは異なります。ロジット方式には出力計算1回の追加があり、両方式のコストは同じではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

ループ型Transformerは、共有するブロックを反復ループの中で繰り返し実行することで、パラメータ効率を実現する。各ループからは同じ次トークンを復号できる中間表現が得られるが、標準的な復号では前の状態を捨てている。早いループほど計算量が少ないため、この反復構造は、補助モデルや外部での学習なしに、対応の取れた弱い予測と強い予測の組を自然に提供する。 私たちは、最終予測をより早い反復パスと対比させてトークン選択を導く、追加学習不要の対比的復号フレームワークLoopCDを提案する。出力計算を1回追加するロジット空間での方式LoopCD-Logitsと、出力計算の追加負担がない隠れ状態空間での方式LoopCD-Hiddenを用意する。4つのループ型Transformer系列にわたり、LoopCDは全反復深さで大きく一貫した改善を示した。LoopCD-LogitsはOuro-2.6B-ThinkingのAIME 2024のpass@1を61.88%から73.33%へ、LoopCD-HiddenはHuginnのHumanEvalのpass@1を22.56%から31.71%へ引き上げた。 重要なのは、こうした性能向上によって、反復ループ数を半分にしても、誘導なしの全深さベースラインと同等以上の性能を維持でき、順伝播のFLOPsを22.5〜48.2%削減できる点である。LoopCDは反復途中の状態を有効な誘導信号へ変換することで、推論計算量を大幅に減らしながら、より良い復号品質を達成する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token selection by contrasting the final prediction with an earlier recurrent pass, operating either in logit space with one extra output pass (LoopCD-Logits) or in hidden-state space with zero output overhead (LoopCD-Hidden). Across four looped Transformer families, LoopCD delivers substantial, consistent gains at full recurrent depth: LoopCD-Logits raises Ouro-2.6B-Thinking's AIME 2024 pass@1 from 61.88% to 73.33%, while LoopCD-Hidden lifts Huginn's HumanEval pass@1 from 22.56% to 31.71%. Crucially, these performance gains enable halving the number of recurrent loops while still matching or exceeding full-depth unguided baselines, reducing forward FLOPs by 22.5% to 48.2%. By transforming intermediate recurrent states into effective guidance signals, LoopCD achieves superior decoding quality while substantially reducing inference compute.

著者のコメント

32 pages, 19 figures

arXiv ID: 2610.02185 / 要約の誤りについて