arXiv論文メモ
新着一覧
cs.CV / cs.LG · 査読状況未確認

動画生成の不自然な動きを注意機構から調べる

Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

Yueyan Li, Haibo Wang, Caixia Yuan, Xiaojie Wang

この論文をやさしく読む

ひとことで言うと

動画生成モデルが不自然な動きを作る原因を、初期の動きの計画と注意の届く範囲から調べています。

何に役立つ?

生成動画の物理的な整合性を改善するために、位置埋め込みの周波数を調整するという具体的な変更を提案しています。

この研究の面白いところ

完成した動画を後から修正するのではなく、ノイズ除去の早い段階で候補位置が固定される仕組みに着目しています。

どこまで分かった?

要旨は実験による改善を述べていますが、対象モデル、評価件数、改善幅は記載していません。すべての物理法則への適合が保証されたという結果ではありません。「初」の位置付けは著者らの主張です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

最先端の動画拡散モデルは優れた見た目の品質を達成しているにもかかわらず、現実の物理法則に反する内容をしばしば生成する。既存の解決策は外部の事前知識や専用データに依存するのに対し、本研究ではモデル内部の仕組みを探索して根本原因を調べる。具体的には、テキストから動画を生成する拡散モデルの「動作計画」過程について初の解釈可能性研究を提示し、ノイズ除去の初期段階で動きの軌跡がどのように形成されるかを明らかにする。「まず形、その後に細部」という知見を踏まえ、交差注意の軌跡パターンと各注意ヘッドの因果的な寄与を組み合わせ、動作計画を駆動する特定の注意ヘッド群を同定する。 さらに自己注意の解析から、回転位置埋め込み(RoPE)が空間的な注意の過度な減衰を引き起こすことを示す。このため初期の候補領域が物理的に不自然な位置へ早期に固定され、隣接フレームで妥当な軌跡が抑えられ、生成の失敗が起きる。この根本的な欠陥に対処するため、ノイズ除去の段階に応じてRoPEの周波数を調整する、軽量なアーキテクチャ変更を提案する。この方法は注意の過度な減衰を軽減し、モデルがより良い候補領域を探索して、物理的に整合した動きを形成するのを助ける。最後に、追加学習を行わない実験と学習を伴う実験の両方により、生成動画の物理的な常識性を高める上で本手法が有効であることを確認する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While existing solutions rely on external priors or specialized data, we investigate the root cause by exploring the internal mechanisms of these models. Specifically, we present the first interpretability study on the ''motion planning'' process of text-to-video diffusion models, revealing how motion trajectories form during early denoising stages. Building upon the ''first shape, then details'' finding, we combine cross-attention trajectory patterns with causal head contributions to identify a specific subset of attention heads driving motion planning. Further, our self-attention analysis shows that Rotary Position Embedding (RoPE) induces excessive spatial attention decay. This causes early candidate regions to prematurely lock into physically implausible positions, suppressing reasonable trajectories in adjacent frames and triggering generation failure modes. To address this fundamental flaw, we propose a lightweight architectural modification that scales the frequency of RoPE across different denoising steps. This strategy reduces excessive attention decay, helping the model explore better candidate regions to establish coherent physical motion. Finally, training-free and training-based experiments confirm the effectiveness of our approach in enhancing the physical commonsense of generated videos.

arXiv ID: 2609.23658 / 要約の誤りについて