外部ツールを使い分ける動画生成エージェントの学習
VideoGen-Agent: Reinforcing Video Generation Agents
この論文をやさしく読む
ひとことで言うと
動画生成AIに、生成前の補助や生成後の確認を行うツールの使い方を学ばせています。一度の生成だけでは満たしにくい指示に、複数段階で対処します。
何に役立つ?
対象の見た目を保つ、出来事を指定順に並べるなど、条件の多い動画生成への応用が考えられます。論文での実証は600プロンプトのベンチマークと人間による比較評価です。
この研究の面白いところ
エージェントを学習し直さず、使う動画生成ツールを更新するだけでも得点が上がっています。方策が特定の生成器の能力向上を取り込めるかを検証している点が特徴です。
どこまで分かった?
19.1は得点の差であり、相対的な改善率ではありません。84.3%は指定された相手との比較での選好率です。あらゆる動画指示で同じ改善が得られるかや、処理費用は要旨からは分かりません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動画生成モデルの近年の進歩により、高忠実度で時間的な一貫性を持つ動画生成が可能になった。しかしこれらのモデルは、専門知識、特定の同一性、物理的整合性、出来事の順序を要求するプロンプトを満たすのに苦労することが多い。本論文は、動画生成のために外部ツールを使うよう、複数課題のエージェント型強化学習で訓練したマルチモーダルエージェントVideoGen-Agentを提示する。エージェントは複数ターンのやり取りを通じて、補強、生成、検証のツールを連携させ、プロンプトと途中の観測を意思決定の手掛かりにする。6課題を対象とし、カテゴリ間で件数を均衡させたデータセットで共通の方策を学習する。教師が生成した軌跡による教師あり微調整でツール使用行動を確立し、その後、強化学習で改善する。カテゴリを考慮した複合報酬が、ツール呼出しの妥当性、課題に適したツール使用、生成動画の品質を評価する。さらに、手順に関する知識、単一・複数対象の同一性維持、物理的整合性、場面構成、複数ショットにわたる時間構造を対象とする、学習から分離した600プロンプトのベンチマークVABenchを導入する。VABench上で、VideoGen-Agentは基となるテキスト動画生成器の56.5点から75.6点へ、19.1ポイント改善した。生成ツールを更新すると、エージェントを追加学習せずに得点はさらに86.1へ上昇した。人間の評価者は、最も強い単体ベースラインとの比較の84.3%で、更新後の構成を好んだ。これらの結果は、動画生成課題をまたいだツール使用の学習を支持し、学習済みエージェントが生成ツールのその後の進歩からも恩恵を受けられることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to use external tools for video generation. The agent coordinates augmentation, generation, and verification tools through multi-turn interactions, using the prompt and intermediate observations to guide its decisions. We train a shared policy on a category-balanced dataset spanning six tasks. Supervised fine-tuning on teacher-generated trajectories establishes tool-use behavior, which is then refined through reinforcement learning. A category-aware hybrid reward evaluates tool-call validity, task-appropriate tool use, and generated video quality. We further introduce VABench, a held-out benchmark of 600 prompts covering procedural knowledge, single- and multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6. Upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons. These results support learning tool use across video-generation tasks and show that the trained agent can benefit from subsequent advances in generation tools.
arXiv ID: 2609.24997 / 要約の誤りについて