画像と言語の計算経路を別々に削ってマルチモーダル推論を高速化
MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
この論文をやさしく読む
ひとことで言うと
画像と言語を扱うモデルで、同じ部品でも入力の種類ごとに不要な計算を分けて削り、入力を読み込む処理を速くする方法です。
何に役立つ?
長い画像・テキスト入力を処理するMLLMで、入力処理の計算量を減らす用途に役立ちます。入力トークン自体を減らす方式とも組み合わせられます。
この研究の面白いところ
アテンションヘッド全体を一括で扱わず、画像・テキスト間の経路ごとに削る点と、その疎な計算を実際に速く動かすカーネルまで作る点が特徴です。
どこまで分かった?
1.6倍などの速度比はプリフィル、すなわち入力処理段階の結果であり、生成全体が同じ倍率で速くなると示したものではありません。99.7%の性能保持は12ベンチマークの平均です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
マルチモーダル大規模言語モデル(MLLM)は、長い画像・テキスト系列を処理する際に、大きな推論コストを要する。既存の演算圧縮法はモダリティ単位の冗長性を利用しているが、アテンションヘッド内部と共通のフィードフォワードネットワーク(FFN)のチャネル内部の計算を、ほぼ一まとまりの単位として扱っており、さらに細かい冗長性は十分に調べられていない。本研究では、冗長性が同じアテンションヘッド内のモダリティ間相互作用経路によって異なり、同じFFNチャネルでも画像入力とテキスト入力の実行によって異なることを見いだす。 これに基づき、モダリティを考慮した幅方向の演算枝刈り(MWOP)を提案する。各層で、画像から画像(V2V)、テキストから画像(T2V)、テキストからテキスト(T2T)のアテンション経路を独立に枝刈りし、画像入力とテキスト入力のFFNチャネルを別々に選択する。一次Taylor基準で枝刈りを導き、アテンションの枝刈りとLoRAによる回復学習の後にFFNの重要度を再評価する。得られた細粒度の疎性を実際の高速化につなげるため、経路の疎性に対応したTritonアテンションカーネルと、画像側FFNのコンパクトな実行も開発する。 MWOPはトークン系列を保ちながらアテンションとFFNの計算を削減するため、トークン圧縮と相補的であり、系列長とトークン当たりの計算量を同時に減らせる。LLaVA-OneVision-7Bでは、MWOP単独でプリフィルを1.6倍高速化し、12ベンチマークにわたって平均性能の99.7%を保持する。代表的な二つのトークン圧縮法と組み合わせると、それぞれのプリフィル高速化率を2.0倍と1.9倍から、2.9倍と2.7倍へさらに高める。Qwen2.5-VL-7Bでの結果も、異なるアーキテクチャへの適用可能性を示す。コードは https://github.com/EIT-NLP/MWOP で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each layer, and separately selects FFN channels for visual and textual inputs. A first-order Taylor criterion guides the pruning process, with FFN importance re-evaluated after attention pruning and LoRA-based recovery training. To translate the resulting fine-grained sparsity into practical acceleration, we further develop path-sparse Triton attention kernels and compact visual-side FFN execution. MWOP preserves the token sequence while reducing attention and FFN computation, making it complementary to token compression and enabling simultaneous reduction of sequence length and per-token computation. On LLaVA-OneVision-7B, MWOP alone achieves a $1.6\times$ prefill speedup with 99.7\% average performance retention across 12 benchmarks. Combined with two representative token compression methods, it further increases their prefill speedups from $2.0\times$ and $1.9\times$ to $2.9\times$ and $2.7\times$, respectively. Results on Qwen2.5-VL-7B further demonstrate its applicability across architectures. The code is available at https://github.com/EIT-NLP/MWOP.
arXiv ID: 2610.01434 / 要約の誤りについて