参照画像や動画の要素を正しく受け継ぐ生成AIの評価基盤
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
この論文をやさしく読む
ひとことで言うと
動画生成AIが参照素材に「なんとなく似せる」だけでなく、動きや見た目など必要な要素を正しい対象に引き継げるかを細かく調べる評価基盤です。
何に役立つ?
複数の参照素材から動画を作るモデルの弱点を分類し、学習データを用意するために役立ちます。34万件の学習サンプルを持つデータセットも提案されています。
この研究の面白いところ
全体の似ている度合いだけでなく、ある参照の動きを別の対象に混ぜていないかなど、要素の分離と対応付けまで評価します。
どこまで分かった?
要旨は課題別の性能差を報告していますが、個別モデルの得点は示していません。データセットを公開することと、その利用で特定のモデルが改善することの実証は区別されます。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
参照から動画を生成するR2V(Reference-to-video)生成は、より一般的で多用途な参照制御へと進化し、包括的なomni R2V生成という新しい枠組みが生まれている。しかし、既存のベンチマークはこうした新しい能力に十分対応していない。テスト例が限られた種類と組み合わせの参照しか含まず、評価手順も主として全体的な参照との整合性を評価するため、参照の各要素が適切に保持され、分離され、意図した対象へ割り当てられているかを見落としている。一方、omni R2Vの学習データの構築には高い費用がかかり、適切な学習資源が不足している。 この不足に対処するため、omni R2Vモデルの評価と学習のためのOmniVBenchとOmni-R2V Datasetを導入する。OmniVBenchは、より広い参照の種類、細粒度の制御課題、豊富な参照の組み合わせへとR2V評価を拡張し、内容、動き、スタイル、構造、物語、複数参照の設定にまたがる7つの課題群と18の細粒度課題を扱う。参照要素に基づく評価として、事例ごとに固有のチェックリスト項目12,172個を導入し、意図した参照要素が忠実に維持され、正しく分離されて対象に結び付けられ、指示に従って適切に実現されているかを評価する。 さらに、多様なR2V課題の産業水準の学習資源を幅広い研究コミュニティに届けるOmni-R2V Datasetを導入する。主に大規模な専門制作の映像コーパスを基に、多様な参照の種類と複数参照の組み合わせを含む、加工済みの学習サンプル34万件で構成される。参照と生成対象の対を構築する課題別の処理パイプラインを開発し、omni R2Vデータ構築のための実用的で拡張可能な手順を提供する。先進的な公開・非公開のR2Vモデルを広く評価した結果、OmniVBenchでは課題群や評価軸によって明確な性能差が見られ、現在のR2Vモデルに残る限界が浮き彫りになった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Reference-to-video (R2V) generation is evolving toward increasingly general and versatile reference control, giving rise to the emerging paradigm of omni R2V generation. However, existing benchmarks fall short of these emerging capabilities: their test cases cover limited reference types and compositions, and their evaluation protocols largely assess holistic reference consistency, overlooking whether reference factors are properly preserved, disentangled, and routed. Meanwhile, the high cost of constructing omni R2V training data makes suitable training resources scarce. To address these gaps, we introduce OmniVBench and the Omni-R2V Dataset for evaluating and training omni R2V models. OmniVBench expands R2V evaluation across broader reference types, fine-grained control tasks, and richer reference compositions, covering 7 task families and 18 fine-grained tasks spanning content, motion, style, structure, narrative, and multi-reference settings. We introduce factor-grounded evaluation with 12,172 case-specific checklist items, assessing whether intended reference factors are faithfully preserved, correctly disentangled and bound to their targets, and properly realized according to the instruction. We further introduce the Omni-R2V Dataset, bringing industrial-grade training resources for diverse R2V tasks to the broader research community. Drawing primarily on a large-scale corpus of professional video footage, it comprises 340K processed training samples spanning diverse reference types and multi-reference compositions. We develop task-specific pipelines for reference-target pair construction, offering a practical and scalable recipe for omni R2V data construction. Extensive evaluation of advanced open- and closed-source R2V models reveals clear performance gaps across task families and evaluation dimensions on OmniVBench, highlighting remaining limitations of current R2V models.
arXiv ID: 2609.22069 / 要約の誤りについて