生成動画の多様性が動きやカメラで失われる場所を調べる
DiVid: Diagnosing Dimension-Specific Diversity Collapse in Video Generation Models
この論文をやさしく読む
ひとことで言うと
生成動画が見た目には多彩でも、動きやカメラワークだけが似通っていないかを、6つの側面に分けて調べる評価方法です。
何に役立つ?
動画生成モデルの弱点が被写体、背景、動きなどのどこにあるかを診断するのに役立ちます。考えられる用途は、弱い側面に対応した学習目標や制御方法の設計です。
この研究の面白いところ
指示から外れた動画を除いてもモデルの順位が残ることを確認しています。また、自由に生成させると定番に戻る問題と、具体的に違いを指定しても実現できない問題を区別しています。
どこまで分かった?
要旨は代表的モデルを評価したと述べますが、モデル数や改善の数値は示していません。診断枠組みの提案であり、動きやカメラの多様性を改善する学習法の効果まで報告したものではありません。枠組みは公開予定とされています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動画生成モデルは目覚ましく進歩しているが、同じプロンプトから繰り返し生成すると非常によく似た出力になりやすく、創作の探索における有用性が制限される。既存の多様性評価は主に全体を表すスカラー指標に依存し、動画の時空間のどこで多様性が失われているかを見えにくくしている。 本研究では、動画生成の多様性を、意味、スタイル、被写体、シーン、動き、カメラという解釈可能な6つの側面へ分解する診断枠組みDiVidを導入する。各側面を再現可能なコンピュータービジョンの処理系で測定し、品質や指示への忠実さと併せて分析することで、潜在的なトレードオフを検討する。代表的な動画生成モデルの体系的な評価から、多様性は側面ごとに大きく異なり、全体の多様性スコアが高いモデルでも、特定の要因、とりわけ動きとカメラでは多様性が失われることが分かった。指示に忠実でない生成結果を除いても、これらの順位は保たれており、プロンプトから外れた出力による見かけの違いではなく、能力の違いであることが示される。 測定に加えて、統制されたプロンプト介入から2つの基本的な障壁を特定する。1つは、自由度の高いプロンプトでモデルが支配的なパターンへ戻る「既定モードへの収束」である。もう1つは、多様な代替案を明示的に要求しても忠実に実現できず、とりわけ時間的な要因で問題となる「実現の隔たり」である。時間的要因で忠実さの低下がより大きいことは、文章だけで動きやカメラの変化を制御する難しさを示している。DiVidは、多様性が存在するかの測定から、どこで、なぜ失われるかの診断へと研究を進め、各側面を考慮した学習目標や制御信号について実行可能な方向性を示す。多様で制御可能な動画生成の今後の研究を促すため、この枠組みを公開する予定である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Despite remarkable progress, video generation models often produce highly similar outputs when repeatedly sampled from the same prompt, limiting their usefulness for creative exploration. Existing diversity evaluations primarily rely on global scalar metrics, which obscure where diversity collapses in the spatiotemporal space of videos. We introduce DiVid, a dimension-level diagnostic framework that decomposes video generation diversity into six interpretable dimensions: Semantic, Style, Subject, Scene, Motion, and Camera. Each dimension is measured through a reproducible computer-vision pipeline and analyzed alongside quality and instruction faithfulness to examine potential trade-offs. Systematic evaluation of representative video generation models reveals that diversity is highly dimension-specific: models with strong global diversity scores still collapse on specific factors, particularly Motion and Camera. These rankings persist after filtering unfaithful generations, indicating genuine capability differences rather than off-prompt outputs. Beyond measurement, controlled prompt interventions identify two fundamental bottlenecks: default mode convergence, where models fall back to dominant patterns under open-ended prompts; and realization gaps, where models fail to faithfully realize diverse, explicitly requested alternatives, particularly for temporal factors. The larger faithfulness losses for temporal factors highlight the difficulty of controlling motion and camera variation through text alone. DiVid thus shifts the study of diversity from measuring whether it exists to diagnosing where and why it collapses, and provides actionable directions for dimension-aware training objectives and control signals. The framework will be released to facilitate future research on diverse and controllable video generation.
arXiv ID: 2610.01661 / 要約の誤りについて