arXiv論文メモ
新着一覧
cs.MM / cs.AI / cs.CL / cs.CV / cs.LG · 掲載先の記載あり

効率的マルチモーダル学習のモデルからシステムまで

From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning

Pan Wang, Siwei Song, Hui Ji, Siqi Cao, Heng Yu, Zhijian Liu, Huanrui Yang, Yingyan Celine Lin, Beidi Chen, Mohit Bansal, Xiaoming Liu, Pengfei Zhou, Ming-Hsuan Yang, Tianlong Chen, Jingtong Hu

この論文をやさしく読む

ひとことで言うと

マルチモーダル学習の効率化を、モデル、アルゴリズム、システムの3層に分け、300以上の研究を整理したサーベイです。

何に役立つ?

効率化の議論をモデル内部だけに閉じず、実行方法やハードウェアまで含めて設計するときの見取り図になります。用途ごとにどの層を最適化するかを考える際の参考になります。

この研究の面白いところ

効率・有用性・プライバシーを別々の問題ではなく、層をまたぐ協調設計のトレードオフとして扱っている点が特徴です。MLLMを例に、研究の進展を構造調整からフルスタックの資源管理まで連続的に捉えています。

どこまで分かった?

サーベイであり、各手法を新たな共通実験で直接比較した研究ではありません。要旨は今後の課題を示していますが、個別の用途でどの設計が最適か、また提案する方向性が実際にどれだけ性能やコストを改善するかは未確定です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

マルチモーダルモデルの急速な拡大によって、計算量、メモリ、配備に関する大きなボトルネックが明らかになり、効率的マルチモーダル学習(EML)は重要な研究フロンティアとなっている。集中的な進展がある一方で、学習スタックのどの部分で、どのように効率が実現されるのかについての一貫した理解は、依然として断片的である。 本サーベイは、モデルからシステムへ至る初の構造化された分類法を導入し、EMLの全体像を体系化する。300を超える代表的研究から知見を抽出し、モデル、アルゴリズム、システムという3つの階層で、それぞれアーキテクチャの簡素化、実行の洗練、ハードウェアを意識したオーケストレーションを扱う。単なるカテゴリー別のレビューにとどまらず、これらの層の垂直的な相乗効果を方法論的に統合し、層をまたぐ協調設計が「効率・有用性・プライバシー」の基本的なトレードオフにどう寄与するかを明らかにする。 さらに、マルチモーダル大規模言語モデル(MLLM)を統合的なケーススタディとして、初期の構造調整から現代のフルスタック資源オーケストレーションに至る分野の発展過程を追跡する。多様な領域に向けた全体的な議論と用途別の最適化設計図を提示し、効率が後付けの制約ではなく、モデルの基本設計から内在的・創発的に生じる性質となる、自己調整型知能へのパラダイムシフトを提唱する。最後に、EML研究の進路を決める未解決課題と今後の方向性を示す。本サーベイは、高性能で汎用的なだけでなく、本来的に効率的で広範な配備に備えたマルチモーダルシステムのための構造化された枠組みを確立する。継続的に更新される版は https://github.com/pwang322/Efficient-Multimodal-Learning-Survey で公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
掲載先の記載あり

著者による掲載先の記載:Transactions on Machine Learning Research, 2026。出版社での独立確認は未実施です。

arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The rapid expansion of multimodal models has surfaced formidable bottlenecks in computation, memory, and deployment, catalyzing the rise of Efficient Multimodal Learning (EML) as a pivotal research frontier. Despite intensive progress, a cohesive understanding of what, how, and where efficiency is manifested across the learning stack remains fragmented. This survey systematizes the EML landscape by introducing the first structured, model-to-system taxonomy. We distill insights from over 300 seminal works into three hierarchical levels--model, algorithm, and system--addressing architectural parsimony, execution refinement, and hardware-aware orchestration, respectively. Moving beyond a purely categorical review, we offer a methodological synthesis of the vertical synergies between these layers, elucidating how cross-layer co-design contributes to the fundamental "Efficiency-Utility-Privacy" trade-off. Through an integrative case study of Multimodal Large Language Models (MLLMs), we trace the field's evolutionary trajectory from initial structural adjustments to modern full-stack resource orchestration. Furthermore, we provide a holistic discussion and application-specific optimization blueprints for diverse domains and posit a paradigm shift toward self-regulating intelligence, where efficiency is an intrinsic, emergent property of the model's fundamental design rather than a post-hoc constraint. Finally, we present open challenges and future directions that will define the trajectory of EML research. This survey establishes a structured framework for multimodal systems that are not only high-performing and generalizable but natively efficient and ready for ubiquitous deployment. A continuously updated version is available at https://github.com/pwang322/Efficient-Multimodal-Learning-Survey.

著者のコメント

TMLR

arXiv ID: 2609.19445 / 要約の誤りについて