arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

動画の説明が本当かを確かめる500本の評価集

VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration

Changbeen Kim, Junwon Chang, Kipyo Kim, Risa Shinoda, Kuniaki Saito, and Donghyun Kim

この論文をやさしく読む

ひとことで言うと

動画AIに説明を作らせるだけでなく、その説明の一文一文が実際の映像にあるかを見抜けるか調べる評価集です。

何に役立つ?

動画理解モデルの選定や改善で、説明生成と誤り検出を別々に評価するのに役立ちます。短い動画から90分の動画まで含み、長さや複雑さによる弱点の違いを調べられます。

この研究の面白いところ

明らかに無関係な文ではなく、AIが生成したもっともらしい誤記述を難しい負例として使っています。説明できる能力と、説明を検証できる能力の両方を問います。

どこまで分かった?

評価対象は500本の動画と人が確認した文単位のラベルです。要旨には各モデルの具体的な成績や誤記述率は記載されておらず、すべての動画分野を網羅するとは述べていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

動画大規模言語モデル(Video-LLM)は近年高い性能を示しているが、動画を細部まで理解する能力を信頼できる形で評価することは依然として難しい。既存のベンチマークは質問応答や正解キャプションとの一致に頼ることが多く、表面的な手掛かりや不完全な注釈によってモデルが成功する場合がある。そこで、詳細な動画キャプションに含まれる各出来事が動画によって裏付けられているかをモデルに検証させるベンチマーク、VidOmni-Benchを導入する。 VidOmni-Benchは、5種類の複雑さと4秒から90分までの多様な長さにわたる500本の動画で構成される。これらの軸に沿って動画を集めた後、多様なVideo-LLMで詳細なキャプションを生成し、人が確認した文単位のラベルを得る。誤った出来事を含む文は、評価における判別困難な負例として用いる。 VidOmni-Benchでの実験から、3つの主な知見が得られた。(1)Video-LLMは詳細な動画キャプションの生成で、事実にない記述を頻繁に作り出す。(2)検証役としても苦戦し、もっともらしいが誤った出来事の記述を確実に見つけられない。(3)モデルの弱点は動画の複雑さと長さによって異なり、現在のVideo-LLMにはモデルごとに異なる多様なボトルネックがある。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-18(UTC)
最新改訂
2026-09-18 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

While Video Large Language Models (Video-LLMs) have recently demonstrated strong performance, reliably evaluating their fine-grained video understanding remains challenging. Existing benchmarks often rely on question answering or ground-truth caption matching, where models may succeed through superficial cues and incomplete annotations. To this end, we introduce VidOmni-Bench, a benchmark that requires models to verify whether each event in dense video captions is supported by the video. VidOmni-Bench consists of 500 videos spanning five complexity types and diverse durations from 4 seconds to 90 minutes. After collecting videos along these axes, we use diverse Video-LLMs to generate dense captions and obtain human-verified sentence-level labels, where sentences containing incorrect events serve as hard negatives for evaluation. Our experiments on VidOmni-Bench reveal three key findings: (i) Video-LLMs frequently generate hallucinated descriptions in dense video captioning; (ii) they also struggle as verifiers, failing to reliably detect plausible but incorrect event descriptions; and (iii) model weaknesses vary across video complexity and duration, revealing diverse, model-specific bottlenecks in current Video-LLMs.

arXiv ID: 2609.21521 / 要約の誤りについて