arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

量子化による機能低下をパープレキシティだけでは捉えられない

Perplexity Cost Understates What Activation Quantisation Breaks

Anish Sathyanarayanan

この論文をやさしく読む

ひとことで言うと

LLMの活性値量子化で、平均的な指標が個別の能力低下を見逃すことを示した研究。

何に役立つ?

量子化後のモデルを評価する際、検索などの能力を別に測る必要性を検討する材料になる。

この研究の面白いところ

誤差の大きさだけでは説明できず、座標との整合を変えると機能が回復した点。

どこまで分かった?

調査は指定したモデル、量子化方式、課題での結果であり、すべての能力の低下を一つの指標で説明したものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

活性値の量子化は通常、モデルが予測する全トークンについて平均したパープレキシティで評価される。本研究は、その平均値から量子化器がどの計算を損なうか分かるのかを調べる。パープレキシティは全体の指標としては有用で、四系列の12モデルにおける780件のモデル内比較では、パープレキシティが低い方式の方が、帰納的な処理と検索能力もより多く維持した例が、それぞれ97.9%と96.0%だった。しかしパープレキシティが1.2~1.5倍に上がる程度でも、帰納的処理の正解率は元の0.959を維持する一方、検索能力は0.554まで下がり、この差は平均値に現れない。差は大きさだけでなく構造にも関係する。同じチャンネルごとの大きさに合わせたガウス雑音では機能がほぼ保たれ、量子化誤差の大きさを固定して符号だけを無作為化しても同様に害が小さかった。誤差の大きさを変えず座標との整合だけを変える回転基底で量子化すると、単一ブロックへの介入でトークン当たり平均3ビットの条件下、帰納的処理は0.001から0.980へ回復した。ただし同条件で検索能力の回復は0.694にとどまった。端から端まで平均4ビットで量子化した場合、帰納的処理は0.968、検索能力は0.534となった。この傾向は最大320億パラメータの別の二モデルでも見られ、試した配備構成では活性値を4ビットにしたAWQでも見られた。パープレキシティの目標値は活性値変換の平均的費用を抑えるが、どの計算が残ったかまでは示さない。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Activation quantisation is usually evaluated with an aggregate metric, perplexity, averaged over every token a model predicts. We ask whether that average identifies which computations a quantiser damages. Perplexity turns out to be a reliable aggregate signal: across 12 models from four families and 780 within-model comparisons, the arm perplexity prefers also retains more induction and more retrieval in all but 2.1 and 4.0 percent of cases respectively. But where perplexity has risen by only a factor of 1.2 to 1.5, induction still keeps 0.959 of its intact accuracy while retrieval has already fallen to 0.554, a gap the aggregate number does not surface. This gap has structure, not just size: a matched Gaussian-noise control of the same per-channel magnitude leaves it largely intact, and randomising only the sign of the quantisation error, every magnitude held fixed, is nearly as harmless, so magnitude alone does not explain the damage. Quantising in a rotated basis, which changes coordinate alignment without changing error magnitude, restores induction from 0.001 to 0.980 at three average bits per token in a single-block intervention, though retrieval recovers less completely at the same setting (0.694); end-to-end at four average bits, induction reaches 0.968 and retrieval 0.534. The pattern holds on two further models up to 32B parameters and, in the deployed configurations we tested, under AWQ once activations are pushed to 4 bits. A perplexity target bounds the average cost of a transformation applied to the activation; it does not, by itself, show which computations survived.

著者のコメント

24 pages (9 main text), 12 figures, 26 tables

arXiv ID: 2609.23125 / 要約の誤りについて