arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

映像内の実情報を優先して異常部分を補修する

Copy What Is Seen, Generate What Is Not: Training-Free Anomaly-Aware Video Restoration

Zhida Qu, Shengchao Chen

この論文をやさしく読む

ひとことで言うと

監視映像の補修で、他フレームに実際に見える背景をまずコピーし、見えない部分だけAIで補う。

何に役立つ?

異常部分を除去した映像を作る際、生成に頼る範囲を狭め、背景の忠実度や時間的な安定性を高める方法として役立つ。

この研究の面白いところ

検出・背景情報・生成・検証をつなぎ、すべてを描き直すのでなく映像に残る証拠を優先する。

どこまで分かった?

三つのデータセットでの評価であり、正解マスクを与えた結果と自前の検出マスクによる結果は別条件である。生成された部分は、元の場面を観測して復元した証拠とは異なる。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

異常を検出する監視システムは、その映像の補修も必要とすることが多い。しかし二つの課題は別々に研究されてきた。追加学習を不要とする異常検出器はスコアやラベルの出力で止まり、追加学習不要の動画編集は検出器ではなくユーザーのプロンプトに応答する。本論文は、固定済みの事前学習モデルだけでこの隔たりを埋め、映像からコピーできる証拠がない場所に限って内容を生成する、異常を考慮した動画補修AVRを提案する。 まず動きの証拠でオープンボキャブラリの候補を選別し、時空間マスクを作る。次に、映像から算出した背景の事前情報で、異常による遮蔽がどこかの時点で外れたすべての画素を埋め、どのフレームにも映っていない部分だけを拡散モデルで合成する。さらに固定済みの検証器が、従来型、事前情報に固定した方式、背景条件付き方式のいずれの補修器を信頼するかを映像ごとに決める。 三つの監視映像データセットで、参照映像を完全に持つ異常の注入と、実際の異常の両方を用いて広範に実験した。AVRは正解マスクを与えた条件でフレーム全体の忠実度が最も高く、編集領域内では学習済みの動画補完器三つに匹敵する。また、自ら生成したマスクを使う条件では、検出後に生成するパイプラインを上回り、異常の残留と、自由な拡散生成によるちらつきの双方を抑える。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A surveillance system that detects an anomaly often has to repair the footage as well, yet the two tasks are studied in isolation: training-free anomaly detectors stop at a score or a label, while training-free video editing answers to a user prompt rather than to a detector. This paper proposes AVR (Anomaly-aware Video Restoration), which closes that gap with frozen pretrained models alone and generates content only where the clip offers no evidence to copy. Motion evidence first gates open-vocabulary proposals into spatio-temporal masks. A background prior computed from the clip then fills every pixel the anomaly ever uncovers, leaving diffusion to synthesize only what no frame showed, and a frozen verifier decides per clip whether to trust a classical, a prior-anchored, or a background-conditioned restorer. Extensive experiments on three surveillance datasets, under both full-reference anomaly injection and real anomalies, show that AVR leads full-frame fidelity under oracle masks, matches three trained video inpainters inside the edited region, and outperforms a detect-then-generate pipeline on the masks it produces itself, while suppressing both the residual anomaly and the flicker of free diffusion.

著者のコメント

10 pages, 9 figures, 7 tables

arXiv ID: 2609.18836 / 要約の誤りについて