文章・画像・動画・音声を共通空間で表すOvis-Embedding
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings
この論文をやさしく読む
ひとことで言うと
文章、画像、動画、音声を同じ表現空間に置いて比較・検索する埋め込みモデル。
何に役立つ?
考えられる用途は異なる種類のデータをまたぐ検索。要旨では複数の埋め込み評価での性能を報告している。
この研究の面白いところ
共通の基盤モデルに加え、データの組み方、難例を重視する学習、埋め込み蒸留、次元圧縮を一体で設計している。
どこまで分かった?
要旨には5つの評価名と最先端という結果が示されるが、個別の得点、比較差、運用条件は記載されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
この報告は、テキスト、画像、動画、音声を自然に統合した、著者らが最先端と位置付ける全モダリティ埋め込みモデル群Ovis-Embeddingを紹介する。モダリティごとに別個の処理系を組み合わせる代わりに、共有のマルチモーダル基盤モデルを用い、異なるモダリティを共通の表現空間に符号化する。 具体的な進展は3つある。第一に、事前学習済みQwen-omniモデルを埋め込みの基盤として採用し、低ランク初期化を伴う対照学習で適応させる。第二に、テキスト、画像、動画、音声、それらが交互に現れるマルチモーダルデータを含む広範で高品質なコーパスを構築する。データ効率を高めるため、同じ種類のソースからサンプリングし、課題が一貫し、バッチ内に有益な負例を含むバッチを作る。第三に、難しい事例を重視するfocal lossと、相補的な専門モデルから細かな類似度構造を移す、類似度に基づく埋め込み蒸留を使う。推論時には低ランクの特徴分解によって、性能低下を最小限に抑えつつ次元数を柔軟に選べるコンパクトな埋め込みを得る。 実証評価では、MMEB-v3、MMEB-v2、MVEB、MAEB、RTEBにおいて、Ovis-Embedding群が最先端の性能を達成したと報告する。これらの結果は、テキスト、画像、動画、音声を横断する有効性を示し、モダリティ間の分断を減らして、任意のモダリティ間検索に使える汎用埋め込みモデルを進める可能性を示している。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In this report, we introduce \textbf{Ovis-Embedding}, a state-of-the-art omni-modal embedding family built on native integration of text, image, video, and audio. Instead of assembling separate modality towers, Ovis-Embedding uses a shared multimodal backbone to encode different modalities in a common representation space. Specifically, we make \textbf{three key advances}: (1) \textbf{native omni-modal initialization}: we adopt a pretrained Qwen-omni model as the embedding backbone and adapt it through contrastive training with low-rank initialization; (2) \textbf{data-centric omni-modal training}: we construct a broad, high-quality corpus spanning text, images, video, audio, and interleaved multimodal data. To improve data efficiency, we introduce homogeneous-source sampling to form task-consistent batches with informative in-batch negatives; and (3) \textbf{embedding-specific training and inference optimization}: we use focal loss to emphasize hard examples and similarity-based Embedding Distillation to transfer fine-grained similarity structure from complementary experts. At inference time, low-rank feature decomposition enables compact embeddings with flexible dimensionality and minimal performance loss. Empirical evaluations show that the \textbf{Ovis-Embedding} family achieves state-of-the-art performance on \textbf{MMEB-v3}, \textbf{MMEB-v2}, \textbf{MVEB}, \textbf{MAEB}, and \textbf{RTEB}, demonstrating its effectiveness across text, image, video, and audio modalities. These results highlight the potential of unified omni-modal training to overcome modality fragmentation and advance universal embedding models for any-to-any retrieval.
arXiv ID: 2609.25165 / 要約の誤りについて