動画生成の動きを軌跡の周波数成分で揃える
MotionSpec: Spectral Trajectory Supervision for Motion-Consistent Video Generation
この論文をやさしく読む
ひとことで言うと
生成動画の動きの軌跡を周波数成分とフレーム間の流れの両方で揃える学習方法を提案した。
何に役立つ?
動作が途切れたり不自然に変化したりする動画生成の改善に役立つ可能性がある。実証されたのは要旨に記された動画生成実験での改善である。
この研究の面白いところ
動きの強さだけでなく、振幅と位相を揃えて時間的な構成も制約し、さらに局所的なフローを整える。
どこまで分かった?
要旨は複数の品質面で改善したと述べるが、データセット、比較手法、改善幅の数値は記載していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
文章から動画を生成する技術の進展により、見た目の忠実度が高い映像を合成できるようになったが、現実的な動きの生成はなお難しい。生成動画では、複雑な動きの間に時間的な不連続、動作の進み方の不整合、構造のゆがみが起こり得る。各フレームが現実的に見えても、動きの推移が不整合または不自然なことがある。通常の生成目標では動きに特化した教師信号が限られ、動きの変化を十分に制約できない。本研究では、スペクトル軌跡一貫性(STC)を中心に据えた動きの教師信号の枠組みMotionSpecを提案する。STCは、基準点に対する密な動きの軌跡を構成し、時間方向のフーリエ変換により動きのスペクトル体積へ変換する。予測した軌跡と目標の軌跡についてスペクトルの振幅と位相を揃え、各時間周波数における動きの強さと、動きの時間的な構成の両方を制約する。この軌跡単位の教師信号を補うため、予測動画と目標動画の連続フレーム間のオプティカルフローを揃える局所フロー一貫性(LFC)も導入し、局所的な動きの遷移を安定させる。実験では、視覚的な忠実度を保ちながら、動きの一貫性、時間的な整合性、もっともらしさを一貫して改善した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent advances in text-to-video generation have enabled high-fidelity visual synthesis, yet realistic motion remains challenging. Generated videos may exhibit temporal discontinuities, inconsistent action progression, and structural distortions during complex movements. Even when individual frames appear realistic, the underlying motion may evolve in inconsistent or implausible ways. Standard generative objectives provide limited motion-specific supervision, leaving motion evolution insufficiently constrained. In this paper, we propose MotionSpec, a motion supervision framework centered on Spectral Trajectory Consistency (STC). STC constructs dense anchor-relative motion trajectories and transforms them into motion spectral volumes via a temporal Fourier transform. By aligning the spectral amplitude and phase of predicted and target trajectories, STC constrains both motion strength across temporal frequencies and the temporal organization of motion. To complement this trajectory-level supervision, we introduce Local Flow Consistency (LFC), which aligns consecutive-frame optical flow between predicted and target videos to stabilize local motion transitions. Experiments demonstrate that MotionSpec consistently improves motion consistency, temporal coherence, and plausibility while preserving visual fidelity.
arXiv ID: 2609.28095 / 要約の誤りについて