arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

文章検索の重みを固定したまま画像や音声も検索する

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal

この論文をやさしく読む

ひとことで言うと

既存の文章検索モデルを変更せず、画像・動画・音声などを同じ検索空間へ追加する方法です。

何に役立つ?

考えられる用途は、既存の文章検索の挙動を保ちながら、複数形式の資料をまとめて検索することです。要旨では文章検索の指標と複数モダリティの比較結果が示されています。

この研究の面白いところ

別の大型教師モデルではなく、固定したモデル自身がキャプションを埋め込んだ結果を教師信号にします。文章側を固定する設計と、追加モダリティ側の学習を分けています。

どこまで分かった?

文章検索を悪化させないという説明は、テキストの重みを同一に保つ設計に基づきます。モダリティ全体の優位性は示された比較範囲の結果で、要旨には個別タスクの全スコアや運用時の速度はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

テキスト埋め込みモデルを新たなモダリティへ拡張すると、通常はテキスト検索の品質が低下する。既存の全モダリティ対応の埋め込みモデルは、数十億規模のパラメータでこれを補っている。本研究では、テキスト側のパラメータを一切更新せずに、テキスト、発話、音響、画像、動画、視覚情報の豊富な文書を、単一の共通コサイン空間へ写す9億パラメータのモデルOmni-Embed-Miniを示す。 中心となる着想は、教師信号のために別の埋め込みモデルを必要としないことである。各メディアサンプルに、段階的に生成した情報密度の高いキャプションを対応させ、固定した基盤モデル自身によるそのキャプションの埋め込みを、教師の目標値としてそのまま使う。教師と生徒は同じ基盤モデルの重みを共有するため、その幾何構造はバイト単位で同一であり、軽量な射影器と、モダリティエンコーダに段階的に導入するLoRAアダプタで整列できる。学習ではMatryoshka SigLIPの対照損失と、エンコーダの改善に伴い負例がより難しくなるオンラインのハイブリッド難負例抽出を組み合わせる。視覚と言語に元から対応した基盤モデルへ置き換えることで、この学習方法を23億パラメータ版にも適用する。 Omni-Embed-Mini-0.9Bはテキストの重みを基盤モデルとビット単位で同一に保つため、学習によってテキスト検索が悪化することはない。MTEB-v2 BEIR-8のnDCG@10は49.57である。その一方で、さらに5つのモダリティに拡張し、比較したすべてのオープンな全モダリティ埋め込みモデルに対し、パラメータ数は約2.7分の1から9.5分の1である。23億パラメータ版は非公開モデルgemini-embedding-2と競合する性能を持ち、全モダリティの平均ではわずかに上回る。モデル、コード、データ、評価用プログラムはプロジェクトページ https://omniembed.cvmbzuai.com で公開している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascaded caption, and the teacher target is simply the frozen backbone's own embedding of that caption. Because teacher and student share the same backbone weights, they inhabit byte-identical geometry, and lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment. Training combines a Matryoshka SigLIP contrastive loss with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves. The recipe carries over to a 2.3B variant by swapping in a native vision-language backbone. Omni-Embed-Mini-0.9B keeps its text weights bit-identical to the backbone, so training cannot regress text retrieval (49.57 nDCG@10 on MTEB-v2 BEIR-8), while extending it to five additional modalities, and is ~2.7x to 9.5x smaller than every open omni embedder we compare against. The 2.3B variant is competitive with the closed gemini-embedding-2, edging ahead of it on the overall-modality average. Models, code, data and evaluation harness are on our project page: https://omniembed.cvmbzuai.com

著者のコメント

Findings of EMNLP 2026. 26 pages, 8 figures, 14 tables. Project page: https://omniembed.cvmbzuai.com

arXiv ID: 2610.02148 / 要約の誤りについて