arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

高解像度でも効率を保つ汎用画像表現モデルIronViT

IronViT: Toward Efficient Generalist Visual Representation Learning

Jiaxi Huang, Yueqi Hu, Xin Zhu, Xiaopeng Zhang, Huiting Qiao, Yanglin Zhang, Zefeng Ji, Rongxue Li, Yifei Xu, Huiying Yu, Wei Liu, Jiayin Zheng, Yinggan Xu, Peipeng Chen, Yin Zhang, Jian Yao

この論文をやさしく読む

ひとことで言うと

さまざまな画像理解能力を先に一つのモデルへ集めてから、高解像度で効率のよい構造へ移す視覚モデル。

何に役立つ?

認識、検索、マルチモーダル理解、ロボット学習など複数の用途に使う視覚エンコーダーを、高解像度でも効率よく動かす手掛かりになる。

この研究の面白いところ

複数の専門モデルから効率的な構造へ一度に蒸留すると品質が落ちたため、softmaxモデルを橋渡しにして段階的に移した。

どこまで分かった?

要旨は評価したタスク群と基盤モデルとの比較を述べるが、具体的な計算量や各タスクの数値は記載されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

汎用の視覚エンコーダーは、意味、空間配置、言語との対応、行動に関わる手掛かりを一つの表現に取り込む必要がある。しかし、現在の高性能な視覚基盤モデルを支えるsoftmax注意機構は、高解像度で計算コストが非常に高くなる。複数の専門モデルの知識を効率的な構造へ直接蒸留するのは自然な解決策だが、本研究では、異なる能力を同時に統合しながら異なるトークン混合構造へ適応させると、表現の質が下がることを見いだした。 そこで、計算を制限する前に能力をまとめるという方針のIronViTを導入する。まず相補的な専門モデルをsoftmax注意機構の能力橋渡しモデルへ蒸留し、その統合表現を、softmax注意と線形注意を組み合わせたエンコーダーへ段階的に移す。専用のデータ処理手順によって、情報密度を高め、対象領域を広げた蒸留用データも選別する。 認識、検索、画素などの密な予測、マルチモーダル理解、ロボット学習にわたり、IronViTは主要な専門・汎用視覚エンコーダーと競合する性能を示した。softmaxの橋渡しモデルは、評価した基盤モデルの中でマルチモーダル理解とロボット学習の総合性能が最も高かった。一方、混合型エンコーダーは幅広い転移性能を保ち、入力解像度が高くなるほど効率面の利点が大きくなった。これらの結果は、構造を変える前に能力を統合すれば、従来のsoftmax注意機構の高解像度での大きなコストを引き継がずに、汎用視覚エンコーダーを得られることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A generalist vision encoder must capture semantic, spatial, language-aligned, and action-relevant cues within a unified representation, yet softmax attention underlying today's most capable visual backbones becomes prohibitively expensive at high resolution. A natural attempt to address both challenges is to distill multiple specialist teachers directly into an efficient architecture. We find that directly coupling these objectives degrades representation quality, as the student must simultaneously reconcile heterogeneous capabilities and adapt them to a different token-mixing architecture. We introduce IronViT, built on a simple principle: consolidate capabilities before constraining computation. IronViT first distills complementary specialists into a softmax attention capability bridge, then progressively transfers the consolidated representation to a hybrid softmax-linear attention encoder. A purpose-built data pipeline further curates the distillation corpus for higher information density and broader domain coverage. Across recognition, retrieval, dense prediction, multimodal understanding, and robotic learning, IronViT is competitive with leading specialist and generalist vision encoders. The softmax bridge achieves the strongest aggregate performance in multimodal understanding and robotic learning among the evaluated backbones, while the hybrid encoder retains broad transfer performance with an efficiency advantage that grows with input resolution. Together, these results show that consolidating capabilities before architectural conversion can yield a generalist visual encoder without inheriting the prohibitive high-resolution cost of conventional softmax attention.

arXiv ID: 2609.29252 / 要約の誤りについて