動画生成のため複数ショットの指示を整えるWanPE
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
この論文をやさしく読む
ひとことで言うと
文章から動画を作る前に、複数のショットを通じて映像の指示を具体化するモデル。
何に役立つ?
動画生成向けの指示文の作成と評価に役立つ。要旨ではWan3.0と組み合わせた人間の選好評価を報告している。
この研究の面白いところ
105万本の動画で学習した3970億パラメータのモデル。約1万1,000件の比較で、30秒の条件では元の指示より選好スコアが50.86ポイント高かった。
どこまで分かった?
効果はWan3.0との組み合わせとWanPEvalの評価条件で示された。ほかの動画生成器や実制作で同じ改善が得られるかは要旨から分からない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動画生成は、映像制作の脚本に相当する文章から始まり、画素として具体化される。現代の動画生成器は30秒までの映像を生成し、複雑な条件に従えるようになっているため、文章の指示が、複数ショットにまたがる動作、カメラの動き、照明、音の計画を大きく左右する。本論文は、実世界の105万本の動画で学習し、監督レベルの映像計画を目指す3970億パラメータの指示拡張モデルWanPEを提示する。WanPEは、動画に基づく逆方向の構築によってショットごとの映像計画を作り、Semantic-Consistency GRPO(SC-GRPO)でショット間と時間経過にわたり利用者の要求を忠実に保つ。この能力を評価するため、5~30秒の長さと異なる細かさの意図を含む、人手で注釈したWanPEvalを整備し、約1万1,000件のブラインドな一対比較評価を行う。Wan3.0の動画生成器で使用すると、元の利用者の指示に比べ、人間の選好スコアが5~15秒では10.66~18.84ポイント、30秒の評価では50.86ポイント上がった。要素を取り除く実験では、逆方向の構築が前向きの書き換えより明確に優れ、SC-GRPOは異なるモデル規模にわたって意味の忠実さを安定して保つことが示された。WanPEは、評価した商用サービスの中で5~15秒の条件では首位となり、30秒ではSeedance 2.5と競争力のある結果だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.
arXiv ID: 2609.30221 / 要約の誤りについて