MoE推論の処理をビット単位でCPUとGPUに分担する
RapidMoE: Exploiting Cross-Asymmetry via Adaptive Residual Offloading for Large-Scale MoE Inference
この論文をやさしく読む
ひとことで言うと
大量の専門処理部分を持つMoEモデルを、CPUとGPUで効率よく動かす方法です。専門処理部分を丸ごと移す代わりに、数値表現のビットや残差に着目して仕事を分担します。
何に役立つ?
考えられる用途は、大規模MoEを異種ハードウェア上で推論する際の転送待ちや計算負荷の軽減です。実験ではデコードとプリフィルの高速化が報告されています。
この研究の面白いところ
モデル側の処理の偏りとハードウェア側の能力差を対応づけ、格納方法・処理経路・実行順序をまとめて設計しています。重要なエキスパートの選択も実行中に調整します。
どこまで分かった?
最大3.5倍と2.1倍は実験で得られた最大値です。要旨には評価したモデル、機器、精度の具体的な値や測定条件が示されていないため、すべての環境で同じ高速化が得られるとは判断できません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
Mixture-of-Experts(MoE)の普及に伴い、異種ハードウェアからなるプラットフォームへの導入需要が高まっている。しかし、この導入は、大規模MoEのアルゴリズム上の要求と、特性の異なるハードウェアとの根本的な不整合を浮き彫りにする。既存のCPU・GPU混合推論システムでは、エキスパートをGPUへ読み込む際にPCIe帯域幅がボトルネックになるか、CPUの計算に大きく依存するため、この不整合を解消できない。その結果、資源利用率が低くなり、パラメーター数の拡大に伴って一定の遅延予算を超えることが避けられなくなる。 本論文では、MoEのルーティングによるアルゴリズム上の負荷の偏りと、異種ハードウェアの物理的な能力差との間にある構造的な対応関係を「Cross-Asymmetry」として特定し、活用する。そのために、大規模MoE推論を効率化する残差オフロードシステムRapidMoEを導入する。残差を分割する枠組みにより、オフロードの単位をエキスパートからビットへ転換する方法を提案する。この転換は、(1)コンパクトで分離された格納を可能にするデータ表現、(2)ハードウェアの能力に合う二つの経路へ計算を分けるルーティング戦略、(3)装置間で記憶処理と計算処理の負荷を均衡させてスケジュールする実行並列性、という三つの主要な側面に及ぶ。さらに、新たな統合多階層重要度調停を用いて、実行時に重要なエキスパートの集合を適応的に調整し、精度と遅延のパレート最適境界を確保する。これらの工夫は、内在するCross-Asymmetryを活用し、アルゴリズムとハードウェアの不整合を根本的に解消する。実験結果では、RapidMoEは最先端のオフロードシステムと比べて、デコードで最大3.5倍、プリフィルで最大2.1倍の高速化を達成した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The widespread adoption of Mixture-of-Experts (MoE) has created a growing need for deployment on heterogeneous platforms. However, it exposes a fundamental mismatch between the algorithmic demands of large-scale MoE and the disparate characteristics of hardware.Existing CPU-GPU hybrid inference systems fail to resolve this as they either encounter PCIe bandwidth bottlenecks when loading experts to GPUs, or rely heavily on CPU computation. Consequently, this leads to low resource utilization and inevitable violations of fixed latency budgets as parameters scale. In this paper, we identify and exploit Cross-Asymmetry--a structural alignment between the algorithmic workload skew of MoE routing and the physical disparity of heterogeneous hardware. To this end, we introduce RapidMoE, a residual offloading system for efficient large-scale MoE inference. We propose how RapidMoE leverages a residual-split framework to enable offloading paradigm shift from expert-level to bit-level, which unfolds across three key dimensions: (1) data representation, enabling compact and decoupled storage; (2) routing strategy, partitioning computation into dual paths aligned with hardware capabilities; (3) execution parallelism, scheduling a balanced storage-compute workload across devices. We further employ a novel Unified Multi-Level Importance Arbitration to adaptively adjust the critical expert set at runtime, ensuring the accuracy-latency Pareto frontier. These innovations exploit inherent cross-asymmetry, fundamentally breaking the algorithm-hardware misalignment. Experimental results show that RapidMoE achieves up to 3.5x speedup in decoding and 2.1x speedup in prefill compared to state-of-the-art (SOTA) offloading systems.
著者のコメント
Accepted by EuroSys 2027. 17 pages
arXiv ID: 2610.01265 / 要約の誤りについて