未編集の動画を検索・要約する評価セットMultiVENT-Raw
MultiVENT-Raw: A Benchmark for Retrieval and Reasoning over Raw Videos
この論文をやさしく読む
ひとことで言うと
説明文や編集の少ない生の動画を探し、出来事を報告する能力を測る多言語の評価セットである。
何に役立つ?
考えられる用途は、動画検索・要約モデルの評価である。要旨は人手の関連性判定と重要事実を含むデータセットを公開したと述べる。
この研究の面白いところ
約12万本、5300時間超の動画を使い、検索と出来事報告の二課題を同じ集合で評価できる。
どこまで分かった?
要旨は有力モデルにも難しいと述べるが、個別の正解率や一般の動画全体への代表性は記していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
オンラインの情報は動画で消費されることが増えている。その多くは、携帯電話、手持ちカメラ、防犯カメラなどで連続撮影し、そのままソーシャルメディアや共有サービスへ投稿された「生の動画」である。制作された動画や編集された素人動画には、台本に沿う発話、画面上の文字、図、メタデータがあり、内容を理解する手掛かりとなる。一方、生の動画には通常こうしたものがなく、情報検索や機械による理解がより難しい。 この分野を進めるため、本研究は多言語の評価セットMultiVENT-Rawを公開する。主に生の動画約12万本、総時間5300時間超に、130件の出来事と222件の出来事中心の質問を対応させている。動画の関連性についての人手による判定と、関連動画から人が抽出した重要な事実も含む。MultiVENT-Rawは、指定された出来事に関係する動画を集合から見つける検索課題と、関連動画を対象読者向けの筋の通った報告にまとめる生成課題の双方を支える。有力な基準モデルで評価したところ、最新のマルチモーダルモデルの一部にとっても、両課題は難しかった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Online information is increasingly consumed in video format. Much of this comes in the form of *raw video*: continuous footage taken on a cell phone, with a hand-held camera, or via CCTV, which is then directly uploaded to social media platforms and content sharing services. Whereas professional or even amateur-edited footage tends to feature scripted speech, chyrons, graphics, and metadata that help contextualize its subject matter, raw video typically contains none of these things, making it a much more challenging medium for information retrieval and machine understanding. To facilitate progress in this domain, we release MultiVENT-Raw, a multilingual collection of nearly 120,000 primarily raw videos (over 5,300 total hours), paired with 130 events and 222 event-centric queries, along with human-annotated video relevance judgments and human-extracted key facts for relevant videos. MultiVENT-Raw supports both a retrieval task---to identify videos in the collection relevant to a query event---and a generation task---to summarize event-related videos into a coherent report for a target user. We benchmark strong baselines on MultiVENT-Raw, showing both tasks to be challenging even for some of the latest multimodal models.
arXiv ID: 2609.28437 / 要約の誤りについて