arXiv論文メモ
新着一覧
cs.LG / cs.CL · 査読状況未確認

低精度の言語モデルでは丸め方を処理箇所ごとに変える

Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2

Yohan Chatelain (1), Pablo de Oliveira Castro (2) ((1) Krembil Centre for Neuroinformatics, CAMH, Toronto, Canada, (2) Universite Paris-Saclay, UVSQ, LI-PaRAD, Versailles, France)

この論文をやさしく読む

ひとことで言うと

丸め誤差をランダム化すると有利な場所と、不利な場所が言語モデル内にあることを調べています。

何に役立つ?

低精度推論の丸め規則を選ぶ根拠になります。報告された改善は小規模なDistilGPT-2での精度エミュレーションによるものです。

この研究の面白いところ

誤差の大きさだけでなく、語彙全体に一様な誤差かどうかがsoftmaxへの影響を左右する点を理論解析と実験で結び付けています。

どこまで分かった?

要旨の主な数値は有効数字部6ビットのDistilGPT-2に関するものです。他の規模のモデルで同じ改善幅になるとは示されておらず、実機の速度や消費電力の改善値も記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

低精度のTransformer推論には、確率的丸め(SR)と最近接丸め(RN)のどちらを使うべきだろうか。その答えは、ネットワークのどの箇所に着目するかによって異なる。本研究では数値形式を固定し、個々の演算箇所の丸め規則だけを変えることで、この効果を切り分ける。任意の精度で実験できるように、可変精度確率的丸め(VPSR)アルゴリズムを用いてPRISMのベクトル化丸めライブラリを任意の仮想精度へ拡張し、丸めの判定がハードウェアの浮動小数点演算で厳密に評価されることを証明する。 処理箇所ごとの利害得失に相補的な洞察を与える二つの解析を行う。第一に、線形射影の確率的な前進誤差の上界から、集約する長さをn、丸め精度の尺度をuとすると、SRの誤差の包絡線はO(√n u)で増大する一方、RNではO(n u)となる。この差は低精度で急速に拡大し、とりわけ長い多層パーセプトロン(MLP)の次元縮小射影で顕著になる。第二に、出力softmaxにおける交差エントロピー損失の変化の期待値を、符号付きドリフト、ドリフトの曲率、Fisher情報で重み付けした分散ペナルティへ二次分解し、二つの箇所で逆の挙動が生じる理由を明らかにする。MLPのノイズは主として一様なロジットのシフトとなり、softmaxはこれに対して不変であるため、SRの分散の影響は大きく割り引かれる。一方、出力ヘッドのノイズは語彙全体で一様ではなく、同じことは成り立たない。 有効数字部をt=6ビットとしたDistilGPT-2の観測結果は理論と一致する。MLPにSRを適用するとパープレキシティは全精度の基準値の1.15倍となり、RNでは2.21倍となる。言語モデルの出力ヘッドでは、SRが非一様な分散を導入するのに対し、決定論的なRNは分散を持たないため、優劣が逆転する。MLP出力をt=6とした混合精度構成で、MLPにSR、出力ヘッドにRNを割り当てると、パープレキシティは全精度の基準値の1.10倍以内に収まり、ビット数をそろえたRNに比べて28%低下する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(\sqrt{n} u)$ in reduction length $n$, versus $O(n u)$ for RN, a gap that widens rapidly at low precision and is most pronounced in the long multilayer perceptron (MLP) down-projection. Second, a second-order decomposition of expected cross-entropy loss change at the output softmax into signed drift, drift curvature, and a Fisher-weighted variance penalty reveals why the two sites behave oppositely: MLP noise is predominantly a uniform logit shift to which softmax is invariant, so SR's variance is largely discounted; head noise is non-uniform across the vocabulary and is not. On DistilGPT-2 at $t=6$ significand bits, observations match theory: SR in the MLP raises perplexity to 1.15x the full-precision reference, versus 2.21x for RN. At the language-model head, the ordering reverses because SR introduces non-uniform variance, whereas deterministic RN carries none. In a mixed-precision configuration (MLP output at $t=6$), assigning SR to the MLP and RN to the head brings perplexity within 1.10x of the full-precision reference, a 28% reduction over matched-bit RN.

著者のコメント

35 pages, 10 figures, 4 tables. Code and evaluation pipeline available at https://github.com/big-data-lab-team/fuzzy-llm and archived on Zenodo at https://doi.org/10.5281/zenodo.23066028

arXiv ID: 2610.01889 / 要約の誤りについて