小さなベクトル群で言語モデルの連合学習の通信を減らす
FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization
この論文をやさしく読む
ひとことで言うと
各端末で言語モデルを微調整するとき、やり取りする更新情報を小さなベクトル群で表して通信量を減らす手法です。
何に役立つ?
通信が制約になる連合学習で、更新量を圧縮しつつ精度を維持するために役立ちます。
この研究の面白いところ
情報を小さくするだけでなく、複数端末の低ランク更新を正しく足し合わせる問題も扱っています。片側だけの更新と両側の更新を交互に使います。
どこまで分かった?
評価はQwen2.5モデルでのパープレキシティと圧縮率です。要旨は具体的なプライバシー攻撃への耐性や、全タスクでの精度維持を実証していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
連合学習(FL)は、プライバシーを保護しながら大規模言語モデル(LLM)を微調整できる一方、膨大な通信負担が重要なボトルネックとなっている。さらに、FLに低ランク適応(LoRA)を適用すると、正確な「積の和」(SoP)の実装と、通信効率のよい「和の積」(PoS)の実装の間に、根本的な集約のジレンマがある。これらの課題に対してFedFitを提案する。 第1に、通信負担を大幅に減らすため、互いに重ならない共有ベクトルバンクによるパラメータ化を導入する。2つの小さく互いに重ならない大域的ベクトルバンクから、高次元のアダプタ行列を再構成する。第2に、集約のジレンマに対処するため、交互最適化のスケジュールを設計する。正確な集約が可能な、切り離された単一バンクの更新と、残差スペクトル集約機構で補正する同時更新を交互に行い、SoPとPoSの対立を解消する。さらに、ブロック単位の量子化とクライアント側の誤差フィードバックを組み合わせ、送信ベクトルを一層圧縮する。 提案アルゴリズムの理論的な収束保証も確立する。Qwen2.5モデルを使った広範な実験では、標準的な連合LoRA手法と同程度のパープレキシティを保ちながら、最大100倍高い圧縮率を達成した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.
arXiv ID: 2610.01537 / 要約の誤りについて