arXiv論文メモ
新着一覧
cs.CV · 査読状況未確認

映像モデルの映画的な表現理解を測るCinematicVQA

CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models

Shuo Xing, Pooja Verlani, Balu Adsumilli, Zhengzhong Tu

この論文をやさしく読む

ひとことで言うと

映像モデルが撮影技法だけでなく、その見せ方が物語で果たす役割を理解できるか測る研究です。

何に役立つ?

動画モデルの評価と、映画表現を学ぶための追加学習に使えます。

この研究の面白いところ

映像の見た目の説明と撮影技法の特定に差があり、段階的推論の指示が多くのモデルで逆効果でした。

どこまで分かった?

要旨は評価したモデル群の傾向を述べますが、個別モデルの得点やデータの規模は示していません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

撮影技術は、構図、照明、カメラ操作による視覚的な物語表現であり、観客が映像をどう受け止め、感情的に関わるかを大きく左右する。大規模視覚言語モデルは動画の質問応答で進歩してきたが、既存のベンチマークは主に個別の技法の識別に注目し、その物語上の効果の理解は十分に測ってこなかった。そこで、技法の認識を超えて映画文法の推論を評価するベンチマークCinematicVQAを導入する。併せて、撮影技法と知覚上の効果、物語上の機能を結び付ける構造化表現Cinematic Scene Graphを導入する。 最新の大規模視覚言語モデルを広く評価したところ、モデルが映像の見た目を説明する成績は、その背後にある技法を特定する成績より一貫して高いという意味理解上の差が見つかった。意外なことに、思考過程を段階的に示すプロンプトは一貫した改善をもたらさず、多くのモデルで成績を下げた。これは、現在のモデルには段階的推論を生かすのに十分な映画分野の知識がないことを示唆する。CinematicVQAの訓練データで追加学習すると、特に物語上の機能と複数段階の推論で、一貫した改善が得られた。CinematicVQAは、モデルの映画的理解を厳密に評価するベンチマークであるとともに、映画表現をより理解する動画モデルの訓練データにもなる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable progress in video question answering, existing benchmarks primarily focus on identifying low-level techniques rather than understanding their storytelling impact. To address this, we introduce CinematicVQA, the first-of-its-kind benchmark for cinematic video understanding that goes beyond technique recognition to evaluate film-grammar reasoning, utilizing our introduced Cinematic Scene Graph (CSG), a structured representation that links filming techniques to their perceptual effects and narrative functions. Through comprehensive evaluation of state-of-the-art LVLMs, we reveal a striking semantic gap: models consistently perform higher on describing visual presentations than on identifying the underlying techniques. Surprisingly, Chain-of-Thought prompting fails to provide consistent gains and degrades performance for most models, suggesting that current LVLMs lack sufficient cinematic domain knowledge to benefit from step-by-step reasoning. Fine-tuning on \textsc{CinematicVQA-train} yields consistent improvements, particularly for narrative function and multi-hop reasoning. Overall, \textsc{CinematicVQA} serves both as a rigorous benchmark for cinematic evaluation in LVLMs and as a practical dataset for training more film-aware video models.

著者のコメント

6 pages

arXiv ID: 2609.28813 / 要約の誤りについて