120万件超の編集データで指示に従う動画編集を学習
VideoX-Qwen: Data-Centric Instruction-Based Video Editing
この論文をやさしく読む
ひとことで言うと
動画への追加や削除などを文章で指示する編集モデルを、大量の編集前後データを作る工程から開発しています。
何に役立つ?
考えられる用途は、動画の指定した部分を編集しながら、他の内容や時間的なつながりを保つことです。
この研究の面白いところ
データ生成、選別、指示の充実、画像と動画を使う段階的訓練を一体として設計しています。100例では11指標中9指標で平均値が最良でした。
どこまで分かった?
89%はデータの自動採用率で、編集成功率ではありません。比較評価は100例で、全指標やあらゆる編集依頼で優位だったという結果ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
汎用的な動画編集の進歩は、大規模な対になった教師データの構築と、動画生成基盤の指示駆動編集への効果的な適応に依存する。動画生成と異なり、動画編集は、無関係な被写体、場面構造、動き、時間的連続性を保ちながら、求められた変換を実行しなければならない。そこで、一般的な指示ベースの動画編集に向けた、データ構築とモデル訓練を統合する枠組みVideoX-Qwenを提示する。拡張可能な生産パイプラインは、専門化した生成モデルと理解モデルを、追加、削除、置換、属性編集の相補的な経路にまとめ、その後に品質選別と指示の充実を行う。120万件を超える方向付き動画編集レコードを生成し、主要なタスク群ごとに40万件超を含み、自動採用率は89%である。得られたコーパスは、元動画・指示・対象動画という統一的なインターフェースを通じ、一般的な編集操作を広く構造的にカバーする。さらに、マルチモーダルの意味的条件付けと、元動画の密な潜在表現による誘導を組み合わせる、統一Qwen-Wan編集モデルを開発する。段階的な画像・動画訓練戦略により、マルチモーダル指示インターフェースを整合させ、動画生成器を元動画に条件付けた編集へ適応させ、選別した高解像度データで出力品質を磨く。UniVideoおよびKling O1との100例の比較では、報告した11指標のうち9指標でVideoX-Qwenが最良の平均値を達成した。対象には、指示への追従、編集品質、内容の保持、構造的・知覚的類似性、動画分布の品質が含まれる。大規模なデータ生産システムと統一訓練枠組みを合わせ、より高い能力を持つ指示駆動動画編集の実用的な基盤を提供する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Progress in general-purpose video editing depends on constructing large-scale paired supervision and effectively adapting video-generation backbones to instruction-driven editing. Unlike video generation, video editing must execute a requested transformation while preserving unrelated subjects, scene structure, motion, and temporal continuity. We present VideoX-Qwen, an integrated data-construction and model-training framework for general instruction-based video editing. Our scalable production pipeline organizes specialized generation and understanding models into complementary routes for addition, removal, replacement, and attribute editing, followed by quality screening and instruction enrichment. It produces more than 1.2 million directional video-editing records, including over 400,000 records in each major task group, with an automatic acceptance rate of 89%. The resulting corpus provides broad and structured coverage of common editing operations through a unified source-instruction-target interface. We further develop a unified Qwen-Wan editor that combines multimodal semantic conditioning with dense source-video latent guidance. A progressive image-video training strategy aligns the multimodal instruction interface, adapts the video generator to source-conditioned editing, and refines output quality with selected high-resolution data. In a 100-example comparison with UniVideo and Kling O1, VideoX-Qwen achieves the best mean result on nine of eleven reported metrics, including instruction following, editing quality, content preservation, structural and perceptual similarity, and video-distribution quality. Together, the large-scale data-production system and unified training framework provide a practical foundation for more capable instruction-driven video editing.
著者のコメント
Technical report
arXiv ID: 2609.26015 / 要約の誤りについて