arXiv論文メモ
新着一覧
cs.CL / cs.LG · 査読状況未確認

量子化モデルの配備前に設定を選ぶための誤差評価法

Predicting Quantization Price for Selecting PTQ Configurations Before Deployment

Junbin Qiu, Jian Mu, Weitong Zhang, Yao Shu

この論文をやさしく読む

ひとことで言うと

学習後量子化のさまざまな設定を、配備前に層出力の誤差とコストで比較する理論的な枠組み。

何に役立つ?

ビット数だけでなく数値形式や量子化粒度も含めて、予算内で設定を選ぶ方法の検討に役立つと考えられる。要旨は選択法の導出を述べている。

この研究の面白いところ

層出力の誤差共分散を下流の曲率で評価し、異なる種類の量子化設定を共通の基準で扱う。従来の固定構造でのビット配分も特殊例として含む。

どこまで分かった?

要旨には実験結果や具体的な精度・コスト改善値が記されていない。実際のモデルでどの程度有効かは、この要旨だけでは判断できない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

重み空間で行う学習後量子化(PTQ)では、量子化済みモデルの出力分布がどの程度ずれるか分かる前に、有限精度の数値形式、粒度、量子化器の種類、変換、ビット数を選ばなければならない。既存のPTQ手法は、再構成誤差、ヘッセ行列に基づく感度、変換の効果、下流の損失など、この劣化を構成する重要な要素を予測する。しかし、これらは通常、量子化の構造を固定した後か、別々の設定群の中で評価される。本研究では、重み空間PTQを、層の出力誤差に価格を付けて配備前に設定を選ぶ問題として定式化する。 採用可能な各層の設定を、配備コストを伴う誤差の発生源として扱う。この設定は層出力の誤差共分散Σ_l(α_l)を生み、量子化前のモデルは下流の曲率を用いて、その共分散をρ̂_l(α_l)=1/2 Tr(Ĥ_l Σ̂_l(α_l))として評価する。この「価格」は、量子化前モデルから量子化モデルへの順方向KLダイバージェンスに由来し、基準モデルではその一次項が相殺される。この見方では、再構成誤差や対角成分だけの評価値は、価格を決める要素を落とした簡略な代用指標になる。一方、有限精度の数値形式、コードブック、粒度、等価変換は、それぞれが生む共分散と必要なコストを通して比較できる候補となる。さらにトレースの簡約から、較正時に使う価格表と、予算制約の下で価格に基づいて設定を選ぶ方法を導く。固定された量子化構造の下でのビット配分は、この枠組みの特殊な場合になる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods predict important pieces of this degradation, including reconstruction error, Hessian sensitivity, transformation effects, and downstream loss, but these pieces are usually scored after fixing the quantization geometry or inside separate configuration families. We formulate weight-space PTQ as pre-deployment configuration selection using priced layer-output error. Each admissible layer configuration is treated as an error generator with a deployment cost, which induces a layer-output error covariance $\boldsymbol{\Sigma}_l(\alpha_l)$, and the full-precision model prices that covariance by downstream curvature, $\widehat{\rho}_l(\alpha_l)=\frac{1}{2}\operatorname{Tr}\left(\widehat{\mathbf{H}}_l\,\widehat{\boldsymbol{\Sigma}}_l(\alpha_l)\right)$. The price follows from full-precision-to-quantized forward KL, whose first-order term cancels at the reference model. It turns reconstruction and diagonal scores into reduced proxies that drop price factors, while finite formats, codebooks, granularities, and equivalent transformations become comparable candidates through the covariances they induce and the costs they pay. A trace reduction then yields a calibration-time price table and a budgeted price-guided selector, making fixed-geometry bit allocation a special case rather than the organizing problem.

arXiv ID: 2609.28270 / 要約の誤りについて