arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

専門分野に対応するMoE内部の部品を重みから選んで追加学習する

Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation

Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari

この論文をやさしく読む

ひとことで言うと

MoEモデルの中から、目的の専門分野に関わる部分を重みだけで見つけ、そこを選んで追加学習する方法です。

何に役立つ?

モデル全体の追加学習費用を抑えながら、数学・医療画像などの特定領域へ適応する用途が考えられます。示された課題では更新パラメータ削減と高速化が報告されています。

この研究の面白いところ

専門家部分の役割をデータによる観測だけで探すのではなく、振り分け重みを語彙へ読み出して特定します。計算の疎性が意味的な分業にもつながる点を利用します。

どこまで分かった?

データ不要なのは専門家部分の特定手順です。対象領域への追加学習までデータ不要という意味ではありません。要旨には比較モデルごとの詳細やすべての条件での速度はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

Mixture-of-Experts(MoE)構造は、各トークンを少数の専門家部分だけへ振り分ける疎な計算によって、モデルの容量を拡大する。本研究では、この疎性によってマルチモーダルMoEの内部に自発的な組織化が生まれるかを調べる。モジュール化するよう明示的に学習していなくても、専門家部分がモダリティや領域をまたいで強い意味的専門性を発達させることが分かった。 この構造に基づき、事前学習済みモデルの重みから直接、領域に特化した専門家部分を特定する、データ不要の手法ExpertLensを導入する。これは振り分け器の重みを意味のある語彙トークンへ復号することで行う。この専門性を使い、対象領域に関係する専門家部分だけを選んでファインチューニングすることで、効率的なマルチモーダル適応を行う。 数学、医療、リモートセンシングの課題では、ExpertLensはモデルの21.7〜47.0%のパラメータだけを更新し、平均4.0倍の学習高速化を達成しつつ、全パラメータのファインチューニングと同等以上の性能を示す。また、適応性能と学習効率の両方でLoRAを上回る。これらの結果は、効率化のために導入した疎性が、効率的な適応に直接利用できる意味的なモジュール性を生み得ることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-specialized experts directly from pretrained model weights by decoding router weights into semantically meaningful vocabulary tokens. We leverage this specialization for efficient multimodal adaptation by selectively fine-tuning experts relevant to a target domain. Across math, medical, and remote sensing tasks, ExpertLens matches or surpasses full fine-tuning while updating only 21.7 - 47.0% of model parameters and achieving a 4.0x average training speedup, and outperforms LoRA in both adaptation performance and training efficiency. These results show that sparsity introduced for efficiency can give rise to semantic modularity that is directly useful for efficient adaptation.

著者のコメント

Project page: https://glab-caltech.github.io/expertlens/

arXiv ID: 2610.02123 / 要約の誤りについて