注意機構の問い合わせとキーを組で考えて重みを削減する
QK-Wanda: Coupling Queries and Keys for Unstructured Pruning
この論文をやさしく読む
ひとことで言うと
注意機構で組になって働く問い合わせとキーの関係を使って、言語モデルから削除する重みを選ぶ方法です。
何に役立つ?
追加の重み学習をせずにモデルを枝刈りする際の基準として使えます。ただし、削減後の文章予測やタスク精度も別に確認する必要があります。
この研究の面白いところ
重み単体ではなく相手の射影の情報まで採点に入れ、問い合わせとキーで削減量を配分します。処理時間の増加は小さい一方、局所誤差の改善が全体性能へ直結しない例も示しています。
どこまで分かった?
評価はQK部分の枝刈りです。Llama-3.1-70Bでは再構成誤差が改善してもパープレキシティが悪化し、全モデルで性能が上がるわけではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Wanda(Sunら、2024)は、問い合わせとキーが内積を通じて相互作用するにもかかわらず、各線形射影の内部で重みを独立に採点して大規模言語モデルを枝刈りする。本研究ではQK-Wandaを提案する。マスクを適用せずRoPEの前で定める再構成目的のもとで、問い合わせとキーの各重みを個別に削除したときの損失に基づき採点する。反対側の射影の情報、つまり問い合わせの重みにはキー、キーの重みには問い合わせの情報をWandaのスコアへ加え、両射影が枝刈り予算を共有できるようにする。閉形式のスコアには勾配も重みの更新も不要である。主要実験で用いた校正条件では、枝刈り全体の時間はWandaに比べ、A100で1.3%、H200で3.1%長いだけである。 TinyLlama、Llama 2、Llama 3、Qwen2.5に属する、5億〜720億パラメータの15モデルで、QKのみの枝刈りを評価する。Wandaと比べ、QK-WandaはQK再構成誤差を、スパース率50%で平均60%、80%で平均45%減らす。下流の性能向上はモデルによって異なる。Llama 2 70Bのスパース率80%では、WikiText-2とC4のパープレキシティがそれぞれ20.3%と13.5%下がり、平均ゼロショット正解率は5.94ポイント上がる。Qwen2.5-72Bも改善するが、Llama-3.1-70Bでは再構成誤差が小さいにもかかわらず、パープレキシティが大きく悪化する。これらの結果は、結合を考慮した枝刈り基準の可能性と、モデル品質を予測する指標としての局所再構成の限界の両方を示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Llama 2, Llama 3, and Qwen2.5, spanning 0.5B-72B parameters. Relative to Wanda, QK-Wanda reduces QK reconstruction error by an average of 60% at 50% sparsity and 45% at 80%. Downstream gains depend on the model. At 80% sparsity on Llama 2 70B, WikiText-2 and C4 perplexity decrease by 20.3% and 13.5%, while mean zero-shot accuracy rises by 5.94 percentage points. Qwen2.5-72B also improves, but Llama-3.1-70B has substantially higher perplexity despite lower reconstruction error. These results show both the promise of coupled pruning criteria and the limits of local reconstruction as a predictor of model quality.
著者のコメント
81 pages, including appendices
arXiv ID: 2610.01554 / 要約の誤りについて