動画の区間をまたいで理由を引き継ぐ学習不要の異常検知
Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding
この論文をやさしく読む
ひとことで言うと
動画の各区間を独立に判定するのではなく、前の区間で考えた理由を次へ引き継ぎながら異常を探す手法です。
何に役立つ?
異常検知専用の追加学習やデータセット固有の注釈に頼りすぎず、動画を評価する用途が考えられます。複数の公開ベンチマークと複数モデルで評価しています。
この研究の面白いところ
激しい動きだけを異常と誤認しないよう、時間をまたぐ判断理由と映像側の特徴を照合します。文章の説明を生成するだけでなく、視覚情報との整合性も利用します。
どこまで分かった?
要旨は競争力のある性能と一貫した改善を報告しますが、最良性能の達成を主張しているわけではありません。誤警報率や具体的なスコア、実運用環境での結果は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
動画異常検知(VAD)は、動画内の異常な出来事が起きる時間区間を特定することを目的とする。既存手法の多くはデータセット固有の学習と整備された注釈に依存しており、未知の事象を含むオープンセットの状況への一般化が制約される。大規模視覚言語モデル(LVLM)に基づく近年のゼロショット手法は、この依存を軽減するが、時間的な連続性や構造化された推論に欠けることが多い。 本研究では、VADを逐次的な認知推論タスクとして捉え直す、完全に学習不要な枠組みCog-VADUを提案する。異常検知の思考連鎖プロンプティング(CoADTP)を導入し、LVLMを動画の各区間にまたがる再帰的な推論の連鎖として展開する。構造化した判断理由を時間とともに伝えることで、モデルは暗黙の時間的記憶を保ち、複雑な異常と動きの大きい正常な活動を頑健に区別できるようになる。 信頼性をさらに高めるため、文章による判断理由を視覚埋め込みに整合させる、モダリティ間の再順位付け段階も設計する。意味的一貫性と時間的整合性を確保し、予測をより精密かつ安定にする。複数の公開VADベンチマークでの広範な実験は、Cog-VADUが競争力のあるゼロショット性能を達成することを示す。さらにモデルをまたぐ評価では、CoADTPがモデルによらず推論ベースの異常検知を一貫して改善し、現実の応用に向けた、解釈可能で一般化可能な異常理解を提供することが示された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 掲載先の記載あり
著者による掲載先の記載:Transactions on Machine Learning Research, August 2026。出版社での独立確認は未実施です。
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but often lack temporal continuity and structured reasoning. We propose Cog-VADU, a fully training-free framework that reformulates VAD as a sequential cognitive reasoning task. Cog-VADU introduces Chain-of- Anomaly Detection Thought Prompting (CoADTP), which unrolls an LVLM into a recurrent reasoning chain across video segments. By propagating structured rationales over time, the model maintains implicit temporal memory, enabling robust discrimination between com- plex anomalies and high-motion normal activities. To improve reliability, we further design a cross-modal re-ranking stage that aligns textual rationales with visual embeddings, enforcing semantic consistency and temporal coherence for refined and stable predictions. Extensive experiments on multiple public VAD benchmarks demonstrate that Cog-VADU achieves competitive zero-shot performance. Moreover, cross-model evaluations show that CoADTP consistently enhances reasoning-based anomaly detection in a model-agnostic manner, pro- viding interpretable and generalizable anomaly understanding for real-world applications.
著者のコメント
Published in Transactions on Machine Learning Research (TMLR), 2026. 39 pages
arXiv ID: 2610.01754 / 要約の誤りについて