arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

低ランク勾配の基底共有で連合学習の通信とメモリーを削減

FedLore: Communication and Memory Efficient Federated Learning via Shared Gradient Low-Rank Projection

Junkang Liu

この論文をやさしく読む

ひとことで言うと

連合学習で各端末が勾配を小さく圧縮するとき、圧縮の座標系をそろえ、ラウンドごとに入れ替える方法です。

何に役立つ?

基盤モデルを分散した端末で共同学習する際の、通信量と最適化用メモリーを抑える設計に役立ちます。

この研究の面白いところ

端末ごとには正確な圧縮でも集約すると良い更新方向にならない問題を、共通基底で解消します。基底を変えることで累積更新の表現力も確保しています。

どこまで分かった?

停留性の理論評価には勾配のカバー条件、滑らかさ、分散、異質性の有界性という仮定があります。実験の性能比較は評価した画像・言語課題とベースラインに関するものです。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

基盤モデルの連合学習は、クライアントのメモリーと通信コストに制約される。LoRAに基づく手法は低ランクのアダプターでこれらのコストを減らすが、固定されたランク予算が適応能力を制限しうる。勾配の低ランク最適化はより柔軟である一方、クライアントが独立に部分空間を選ぶと、本研究で「部分空間の断片化」と呼ぶ問題が生じる。ローカルな射影がデータの異質性と相互作用して集約方向に偏りをもたらし、集約によって更新のランクと通信コストが増えることもある。したがって、各クライアントで勾配を正確に圧縮できても、大域的な降下方向が保たれるとは限らない。 本研究では、各ラウンド内で低ランク最適化の基底を共有し、ラウンド間で更新するFedLoreを提案する。共有基底によって低ランク座標での厳密な集約が可能になり、特定した射影の偏りを除去できる。部分空間を更新することで、累積したモデル更新は1ラウンド当たりのランク予算を超えられる。集約の偏りを特徴付け、大域勾配をカバーする条件、標準的な滑らかさ・分散の仮定、および有界な勾配異質性の下で、射影SGD版に対するO(T⁻¹ᐟ²)の停留性評価を確立する。 連合事前学習を含む画像と言語の課題での実験により、FedLoreは、評価した低ランクアダプターのベースラインを上回り、通信と最適化状態のメモリーを減らしながら、全パラメータ学習と同等以上の性能を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Federated training of foundation models is constrained by client memory and communication costs. LoRA-based methods reduce these costs through low-rank adapters, but their fixed rank budget can limit adaptation. Gradient low-rank optimization offers greater flexibility, yet independently chosen client subspaces create a problem we term \emph{subspace fragmentation}: local projections interact with data heterogeneity to bias aggregated directions, while aggregation can increase update rank and communication cost. Thus, accurate local gradient compression need not preserve global descent. We propose \texttt{FedLore}, which shares a low-rank optimization basis within each round and refreshes it across rounds. The shared basis enables exact aggregation in low-rank coordinates and eliminates the identified projection bias. Subspace refresh allows the accumulated model update to exceed the per-round rank budget. We characterize the aggregation bias and establish an $O(T^{-1/2})$ stationarity bound for the projected-SGD variant under a global-gradient coverage condition and standard smoothness and variance assumptions, with bounded gradient heterogeneity. Experiments on vision and language tasks, including federated pre-training, show that \texttt{FedLore} outperforms the evaluated low-rank adapter baselines and matches or exceeds full-parameter training, while reducing communication and optimizer-state memory.

arXiv ID: 2610.01620 / 要約の誤りについて