arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

動画生成モデルの反応からAI生成動画を検出するTRACE

TRACE: Trajectory Representation and Consistency Estimation for AI-Generated Video Detection

Huangsen Cao, Hongkang chu, Siyao Yu, Xin Ding, Jianfeng Dong, Yongwei Wang

この論文をやさしく読む

ひとことで言うと

動画生成モデルが実写とAI生成動画に示す反応の違いを使い、未知の生成器の動画も検出する手法。

何に役立つ?

生成器が変わっても使える合成動画検出器を検討する際の手掛かりになる。

この研究の面白いところ

見た目の不自然さではなく、生成モデル内の速度応答と隣接フレームの一貫性を検出の特徴に使う。

どこまで分かった?

要旨はAIGVDBenchでの改善を述べるが、誤検出率や実環境での性能の具体的な数値は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

動画生成モデルの進歩により、見た目が現実的な映像を合成でき、生成動画の検出が難しくなっている。既存の検出器は外観上の不自然さ、意味の不整合、時間的なパターンに頼ることが多く、これらは特定の生成器に固有で、未知の合成モデルへの汎化を妨げる可能性がある。本研究は、事前学習済みの生成モデルに動画を入力したときの反応が、より転用可能な鑑識上の手掛かりになるかを調べる。実写とAI生成の動画は、事前学習済みFlow Matching動画モデルの下で異なる速度応答を示すという観察が中心である。この違いは、探査に使う事前学習済み動画生成モデルの基盤を変えても残り、外観上の不自然さ以外に転用可能な信号となることを示唆する。そこで、生成過程を考慮したAI生成動画検出の枠組みTRACEを提案する。事前学習済み動画DiTを速度場の探査器として用い、複数のフロー時点で表現を取り出し、隣接フレームの速度差からフレーム間の一貫性をモデル化する。さらに、生成器に依存しにくい表現の学習を促す「実写中心の軌道最適化」という目的関数を導入する。AIGVDBenchでの幅広い実験では、多様な生成器に対して有効に汎化し、学習時に見ていない公開・非公開の動画生成モデルに対し、従来の最良手法を大幅に上回った。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent advances in generative video models have enabled the synthesis of visually realistic content, posing significant challenges to synthetic video detection. Existing detectors often rely on appearance artifacts, semantic inconsistencies, and temporal patterns that may be generator-specific, limitating generalization to unseen synthesis models. We investigate whether responses to a pretrained generative model provide more transferable forensic cues. Our key observation is that real and AI-generated videos exhibit distinct \emph{velocity responses} under a pretrained Flow Matching video model. This distinction persists when different pretrained video-generation backbones are used as probes, suggesting that velocity responses offer transferable forensic signals beyond visual artificts. Motivated by this observation, we propose \textbf{TRACE} (\emph{\underline{T}rajectory \underline{R}epresentation \underline{a}nd \underline{C}onsistency \underline{E}stimation}), a generation-process-aware framework for AI-generated video detection. TRACE leverages a pretrained video DiT as a velocity-field probe to extract representations at multiple flow time points, and models cross-frame consistency through velocity differences between adjacent frames. We further introduce a \emph{Real-Centered Trajectory Optimization} objective that encourages generator-invariant representation learning. Extensive experiments on AIGVDBench demonstrate that TRACE generalizes effectively across diverse generators, substantially outperforming prior state-of-the-art methods on unseen open- and closed-source video generation models.

arXiv ID: 2609.25775 / 要約の誤りについて