生成動画の文字の正確さを300の指示で評価する
VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
この論文をやさしく読む
ひとことで言うと
生成動画がきれいに見えるかに加え、看板や説明などの文字を正しく描けているかを調べるベンチマークです。
何に役立つ?
文字が情報の中心になる動画生成で、映像品質とは別に文字の忠実性を比較できます。広告や科学動画を含む場面が用意されています。
この研究の面白いところ
文字の正確さと場面・動作の要件を分けて測るだけでなく、画像と動画を作って視覚評価で修正するエージェント型の方法も提示しています。
どこまで分かった?
評価は5種類の場面の300プロンプトと11モデルです。最良モデルのWER 0.250は、動画の25%が失敗したという意味ではありません。提案した改善枠組みの具体的な改善量は要旨にありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
近年の動画生成モデルは、自然言語の指示から非常に現実的な動画を生成でき、映像の品質は映画並みに近づいている。しかし、既存の評価ベンチマークは主に視覚的品質、美しさ、物理的なもっともらしさを評価し、日常の場面で情報を伝える重要な媒体である文字には十分な注意を払っていない。生成動画は魅力的に見え、被写体が本物らしくても、場面内の文字を誤って描画することがある。 この見過ごされてきた側面に対処するため、動画生成モデルの視覚的な文字描画能力を体系的に評価するVTR-Benchを導入する。VTR-Benchは、広告や科学動画などの具体的な応用場面に文字を配置し、5種類の場面にわたる、入念に作成した300件のプロンプトを備える。人間の評価との整合を図った自動評価処理を開発し、文字を載せる媒体に応じた書き起こしによって文字の忠実性を、プロンプトごとの一連の問いによって場面と動作の要件を、それぞれ分けて評価する。 評価に加え、キーフレームで誘導するエージェント型の枠組みを導入する。Directorエージェントが画像生成、動画生成、視覚評価を調整し、視覚的なフィードバックを使って反復的な改善と候補選択を導く。11の最先端モデルでの実験では、場面内の文字を正確に描画することに広範な困難が見られ、最良モデルでも全体の単語誤り率(WER)は0.250だった。さらに文字描画の失敗を分析し、現在の動画生成モデルが直面する課題の特徴を明らかにする。これらの結果は、視覚的な文字描画が動画生成の主要な課題であることを示し、改善に向けた実用的な道筋を提示する。コードは https://github.com/hardenyu21/VTR-Bench で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an essential medium for conveying information in everyday scenes. A generated video may appear visually compelling and feature lifelike subjects, yet still render the text within the scene incorrectly. To address this overlooked dimension, we introduce \textbf{VTR-Bench}, a systematic benchmark for evaluating the \textbf{V}isual \textbf{T}ext \textbf{R}endering capabilities of video generation models. VTR-Bench situates text within concrete application scenarios, such as advertisements and scientific videos, with 300 carefully constructed prompts spanning five scenario categories. We develop an automated evaluation pipeline with human alignments that separately assesses text fidelity through carrier-specific transcription and scene and motion requirements through a prompt-specific chain of query. Beyond evaluation, we introduce a \textbf{Keyframe-Guided Agentic Framework} in which a Director agent coordinates image and video generation with visual evaluation, guiding iterative refinement and candidate selection through visual feedback. Experiments on 11 state-of-the-art models reveal widespread difficulties in accurately rendering scene text, with the best-performing model recording an overall word error rate (WER) of 0.250. We further analyze text rendering failures to characterize the challenges faced by current video generation models. These findings highlight visual text rendering as a key challenge for video generation and demonstrate a practical path toward improvement. Code is available at https://github.com/hardenyu21/VTR-Bench.
arXiv ID: 2610.01499 / 要約の誤りについて