arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

構造の違うニューラル網へ学習済み重みを移す

CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

Adir Dayan, Yam Eitan, Haggai Maron

この論文をやさしく読む

ひとことで言うと

学習済みの大きなモデルから、小さなモデルの良い初期重みを予測し、知識蒸留を速くする方法です。

何に役立つ?

モデル圧縮の際に小型モデルを最初から学習する負担を減らす用途が考えられます。要旨では複数の画像処理モデルとINRで蒸留時間の短縮を評価しています。

この研究の面白いところ

元モデルだけでなく変換先の初期値も入力にすることで、構造が違う二つのネットワークのニューロン置換を整合的に扱っています。

どこまで分かった?

普遍性の証明には一般位置とコンパクト集合という条件があります。最大8.89倍は評価内の最大値で、すべてのモデルや圧縮率で同じ高速化を示したという意味ではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

重み空間ネットワークは、他のニューラルネットワークのパラメータを直接扱い、モデルの性質の予測、学習済みモデルの編集、重みの生成などを可能にする。ニューロンの置換など、重み空間に存在する対称性のため、同変性が重要な設計原理となる。しかし、既存の同変な重み空間アーキテクチャは、主としてネットワーク構造を維持する変換について研究されてきた。一方、モデル圧縮や大規模化などの実用的な変換の多くは、学習済みの元ネットワークを異なる構造の変換先ネットワークへ写す。この場合、元と変換先の置換対称性は異なるパラメータ空間に作用するため、同変性の定式化は単純ではない。 この不一致に対処する中心的な着想は、構造間の変換演算子を、学習済みの元ネットワークと変換先ネットワークの初期値という二つの入力を持つものとして定式化し直すことである。これにより、元ネットワークの情報を使って変換先の初期値を改善し、元ネットワークの置換には不変、変換先ネットワークの置換には同変な、構造間変換演算子を定義できる。この定式化に基づき、対称性を保つネットワーク間メッセージ伝播によって両ネットワークを同時に処理するグラフメタネットワークCrossGMNを導入する。一般位置の仮定の下で、CrossGMNがコンパクト集合上の連続な構造間変換演算子に対して普遍性を持つことを証明する。 モデル圧縮を対象にCrossGMNを評価し、小さなネットワークのパラメータを予測することで、その後の知識蒸留を高速化する。2次元・3次元の陰的ニューラル表現(INR)と、MLP、CNN、Vision Transformerを用いた画像分類において、CrossGMNは蒸留を最大8.89倍高速化し、再学習なしで別のデータセットにも転用できた(3.78倍の高速化)。また、単一のモデルで、異種の元アーキテクチャから共通の変換先アーキテクチャへの圧縮を高速化できる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Weight-space networks operate directly on parameters of other neural networks, enabling tasks such as predicting model properties, editing trained models, and generating weights. Weight-space symmetries such as neuron permutations make equivariance a key design principle. However, existing equivariant weight-space architectures have primarily been studied for transformations that preserve the network architecture. In contrast, many practical transformations, including model compression and upscaling, map a trained source network into a target network with a different architecture. In this setting, the source and target permutation symmetries act on different parameter spaces, making equivariance less straightforward to formulate. Our key idea for addressing this mismatch is to reformulate cross-architecture operators with two inputs: a trained source network and an initialization of the target network. This lets us define equivariant cross-architecture operators that refine the initialization of the target network using information from the source network, while being invariant to source-network permutations and equivariant to target-network permutations. Based on this formulation, we introduce CrossGMN, a graph metanetwork that jointly processes both networks through symmetry-preserving cross-network message passing. We prove CrossGMN is universal for continuous cross-architecture operators on compact sets under a general-position assumption. We evaluate CrossGMN for model compression, predicting a smaller network's parameters to accelerate subsequent knowledge distillation. Across 2-D and 3-D INRs and image classification with MLPs, CNNs, and Vision Transformers, CrossGMN speeds up distillation by up to 8.89x, transfers across datasets without retraining (3.78x), and a single model can accelerate compression from heterogeneous source architectures into a common target architecture.

arXiv ID: 2610.01649 / 要約の誤りについて