arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

視覚言語モデルの低ビット量子化で重要な方向を保つ

RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models

Zhiping Wu, Dongdong Ren, Yangchengyu Zhou, Zhengjie Zhang, Wenbin Li, Hongbing Pan, Yang Gao

この論文をやさしく読む

ひとことで言うと

視覚と言語で量子化に敏感な方向を考慮し、VLMを低ビット化する方法。

何に役立つ?

メモリの限られた環境で視覚言語モデルを使うための量子化設計に役立つ。

この研究の面白いところ

リーマン計量で敏感な方向を調べた後、既存の量子化法が使えるユークリッド形式へ変換する。

どこまで分かった?

精度改善はW2A8・W3A8などのベンチマーク条件での最大値として報告される。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模な視覚言語モデル(VLM)は、学習後の量子化によって、厳しいメモリと遅延の制約下でも効率よく使える。しかし多くの量子化法は単一の言語モダリティを持つ大規模言語モデル向けである。ユークリッド幾何の仮定の下で量子化誤差を方向によらない摂動として扱うため、VLMでどの方向が量子化に敏感かを十分に示せない。このため、単一モダリティ向けの方法を直接転用したり、モダリティ別の尺度調整だけを行ったりすると、低ビット条件でビット幅の配分が偏り、性能が不安定になることがある。そこでリーマン幾何に敏感な量子化RGSQを提案し、統一されたFisher–Riemann計量の下で量子化を再構成問題として定式化する。モダリティごとに分けた経験的Fisher因子から作るリーマン多様体の写像によって、各モダリティで敏感な方向を特定し、それらをモダリティを考慮したKronecker構造の計量に融合する。続いて幾何に合わせた回転で局所接空間の座標系を変え、低ビット化による摂動を損失に鈍感な軸へ向ける。最後に白色化変換でリーマン計量の目的関数を等価なユークリッド形式へ移し、既存の単一モダリティ用量子化法が元の仮定のまま複数モダリティの誤差を評価できるようにする。多様な主要VLMベンチマークでは、極めて低ビットのW2A8とW3A8で最高の精度と安定性を示し、VLM向けのMBQやMQuantなどを最大5.9%、単一モダリティへの改善を最大8.6%上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large vision-language models (VLMs) can be efficiently deployed under stringent memory and latency constraints through post training quantization (PTQ). However, most PTQ methods are designed for unimodal large language models (LLMs). These methods treat quantization errors as isotropic perturbations under the Euclidean assumption, which provides weak guidance on directions most sensitive to quantization in VLMs. Consequently, directly adapting unimodal PTQ approaches or solely employing modality-specific scaling often leads to uneven bit-width distribution and inconsistent performance in low-bit settings. To address these challenges, we propose Riemannian Geometry-Sensitive Quantization (RGSQ), which formulates quantization as a reconstruction problem under a unified Fisher-Riemannian metric. RGSQ identifies modality-specific sensitive directions via Riemannian manifold mappings built from modality-partitioned empirical Fisher factors and fused into a modality-aware Kronecker-structured metric. We then apply geometry-aligned rotations to reorient the local tangent frame, steering low-bit perturbations toward loss-insensitive axes. Finally, we apply a whitening transformation that maps the Riemannian objective to an equivalent Euclidean form, enabling standard unimodal PTQ methods to evaluate multimodal quantization error under their original assumptions. Across an extensive and diverse set of mainstream VLM benchmarks, RGSQ achieves the highest accuracy and stability under extremely low-bit settings (W2A8 and W3A8). It outperforms VLM-aware baselines, such as MBQ and MQuant, by up to 5.9% and surpasses single-modality improvements by up to 8.6%.

arXiv ID: 2609.25492 / 要約の誤りについて