arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

層ごとの特徴を保ちながらTransformerの重みを圧縮

GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression

Baher Mohammad, Ammar Ali, Stamatios Lefkimmiatis

この論文をやさしく読む

ひとことで言うと

Transformerの複数層に共通する重みの構造を利用し、追加学習なしでモデルを圧縮する方法。

何に役立つ?

複数層をまとめて圧縮する際、各層の特徴の違いを残しながら重みを減らす設計に役立つと考えられる。

この研究の面白いところ

隣接層を単純に一組にせず、構造的に相性のよい射影を選び、層ごとの較正時の幾何構造を保つ。

どこまで分かった?

要旨は幅広いモデルでの優位を主張するが、圧縮率や精度の具体値は示していない。追加学習不要という条件での結果である。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Transformerには層をまたぐ重複があるが、学習後の圧縮処理は通常、層を個別に最適化するか、層固有の活性化の幾何構造を考慮しない経験則によるグループ分けに頼る。本論文は、層をまたぐ重みの組合せと共有辞書による因子分解を順次最適化する、追加学習不要の原理的な枠組みGeoPairを提案する。隣接層の重みに無理に同じ基底を共有させたり活性化統計を経験則で統合したりせず、構造的に適合する射影を見つけ、各層に固有の較正時の幾何構造をよりよく保つ共有表現を学ぶ。これを構造的な疎性と組み合わせると、機能上の忠実度を犠牲にせず、重みを効率的に分解できる。さまざまな構成、規模、モダリティで最先端の結果を達成し、独立した構造化重み分解や経験則で重みを組にする別方式を一貫して上回ったと報告する。経験則による設計を、収束性のある最適化主導の手順に置き換えることで、異なるモダリティのTransformerを拡張可能に圧縮する理論的基盤を示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.

arXiv ID: 2609.25963 / 要約の誤りについて