arXiv論文メモ
新着一覧
cs.RO / cs.AI · 査読状況未確認

視覚・言語・動作モデルの量子化を閉ループで評価

VLAQuantBench: Closed-Loop Evaluation of Post-Training Quantization for Vision-Language-Action Models

Jiuyi Xu, Qing Jin, Meida Chen, Song Wang, Yang Sui, Yangming Shi

この論文をやさしく読む

ひとことで言うと

ロボット向けVLAモデルを省メモリ化する量子化について、どの層をどの精度にするかで成功率が大きく変わることを調べました。

何に役立つ?

考えられる用途は、VLAモデルを小さいメモリで動かす際に、量子化する層と保護する層を具体的に選ぶことです。

この研究の面白いところ

ある設定では対象層を126から167へ広げるだけで成功率が7.0%から70.5%に変わり、別のモデルでは小さな出力射影一つの保護が重要でした。

どこまで分かった?

主な規模は四モデル、409回の実行、94,574回のシミュレーションです。結果は量子化方法とモデルに依存し、普遍的な層ごとの感度則とはしていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

学習後の量子化は視覚・言語・動作(VLA)モデルのメモリ使用量を減らすが、精度の選択には対象とする層の範囲、数値形式、較正の相互作用を考える必要がある。本研究はVLAQuantBenchという制御された評価を導入し、409回の実行と94,574回のシミュレーションを行った。LIBERO上で四つのモデルを評価し、X-VLAはさらに三種類のシミュレーション・ベンチマーク群でも評価した。 較正せずにW4A4の最近傍丸め量子化を使う場合、π₀.₅モデルの動作出力部で対象層を126から167へ広げると、成功率は7.0%から70.5%へ上がった。観測を固定した再生でも、対応する数値的な回復を確認した。二エピソードを使う較正は、試験した部分集合で重なって起きる深刻な失敗を解消した。一方、同じ平滑化とクリッピングの方法はπ₀の成功率を下げ、OpenVLA-OFTの一連の処理の性能も回復しなかった。OpenVLA-OFTでは、28,672パラメータの出力射影一つを量子化から保護すると、基準に近い成功率が回復した。残る441の対象線形層は、LIBERO-Longで重み3ビット、または四つの課題群すべてで活性値8ビットを維持できた。課題ごとにまとめた区間評価も、大きな失敗と回復の差を支持した。 これらの結果は、すべてのモデルに共通する層の感度規則ではなく、手法ごとの相互作用と具体的な精度割り当てを示す。実際の計算カーネルと物理ロボットでの測定も、正確さの分析を補完する。コード、設定、エピソードの記録は公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-21(UTC)
最新改訂
2026-09-21 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Post-training quantization reduces the memory requirements of vision-language-action (VLA) models, but precision selection must account for the interaction between layer scope, numerical format, and calibration. We introduce \textbf{VLAQuantBench}, a controlled evaluation with 409 runs and 94,574 simulation episodes: four models on LIBERO, with X-VLA additionally evaluated on three simulation benchmark families. Under uncalibrated W4A4 round-to-nearest quantization, expanding a $\pi_{0.5}$ action-head subset from 126 to 167 layers raises success from 7.0\% to 70.5\%. Fixed-observation replay confirms a corresponding numerical recovery. Two-episode calibration removes the severe joint failures in the tested subsets, whereas the same smoothing-and-clipping recipe lowers $\pi_0$ success and does not recover OpenVLA-OFT end-to-end. For OpenVLA-OFT, protecting one 28,672-parameter output projection instead restores near-baseline success: the remaining 441 eligible linear layers retain W3 on LIBERO-Long or eight-bit activations across all four suites. Task-clustered intervals support the large failure and recovery contrasts. These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules. Real-kernel and physical-robot measurements complement the accuracy analysis. Code, configurations, and episode records are publicly available at https://github.com/jiuyixu25/VLAQuantBench.

著者のコメント

28 pages, 35 tables, 4 figures

arXiv ID: 2609.25376 / 要約の誤りについて