音声・映像モデルの時間理解の弱点を調べるAVTrace
AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models
この論文をやさしく読む
ひとことで言うと
映像と音声を扱うモデルが、出来事の時刻、順番、同期を本当に理解しているかを分けて測る評価セットです。
何に役立つ?
映像の内容を説明できることと、時間位置を正しく特定できることを区別してモデルを改善するために役立ちます。
この研究の面白いところ
5つの公開モデルは同期確認で多数派ラベル基準0.556を下回りました。時間に特化した追加学習で一部は改善し、文章の意味の重なりを時間理解の代用にできないことを示しています。
どこまで分かった?
データはsilver-standardで、モデルごとの入力設定も異なります。摂動実験は感度を示しますが原因を切り分けておらず、外部画像課題では追加学習後に悪化した指標もあります。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声・映像などを扱うオムニモデルは動画の内容を説明できる。しかし、出来事を時間上に位置付け、順序を保ち、音声と映像の同期を判断できるだろうか。本研究は、AVTrace(Audio-Visual Temporal Reasoning Assessment and Capability Evaluation)を導入する。これは、開始時点と区間の特定、同期、次の段階の予測、モダリティをまたぐ位置特定、出来事の連鎖の解析、出来事を条件とする理解を対象とした、シルバースタンダードの診断用評価群である。学習例34,114件に加え、カテゴリを均等にした開発用3,500件、テスト用7,000件を含む。 公開された5つのオムニモデルを、それぞれの入力構成で評価する。参照正解を見ずに回答を正規化した後、決定的な規則で採点する。既成の5システムすべてが、同期検証ではテスト集合の多数派ラベル基準値0.556を下回り、連鎖の解析、出来事を条件とした位置特定と理解でも低いスコアとなる。開発集合への摂動では、Qwen3-Omni-30Bがモダリティの除去や視覚入力処理の変更に対して示す感度が課題ごとに異なることが分かるが、その根本原因を切り分けるには至らない。 パラメータ効率のよい時間理解の事後学習は、Gemma4-E4B-itを複数のベンチマーク指標で改善する。外部の画像ベンチマーク3つでは、一部の低下を含めて課題指標の変化は小さい一方、教師強制によるパープレキシティは低下する。これらの結果は、参照文との意味的な重なりを時間的位置特定の代用指標として扱うべきではないこと、またAVTraceが課題固有の弱点を特定し、時間理解の事後学習の試験基盤を提供できることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training examples and category-balanced development and test splits of 3,500 and 7,000 examples. We evaluate five open omni models under their respective input configurations using reference-blind response normalization followed by deterministic scoring. All five off-the-shelf systems score below the test split's majority-label baseline of 0.556 on synchronization verification, and obtain low scores on chain parsing and event-conditioned grounding and comprehension. Development-set perturbations reveal task-dependent sensitivity in Qwen3-Omni-30B to modality removal and changes in visual input processing, without isolating their underlying causes. Parameter-efficient temporal post-training improves Gemma4-E4B-it on several benchmark metrics. On three external image benchmarks, task metrics change modestly, including some degradations, while teacher-forcing perplexity decreases. Together, these findings show that semantic reference-text overlap should not be treated as a proxy for temporal localization, and that AVTrace can identify task-specific weaknesses while providing a testbed for temporal post-training.
arXiv ID: 2609.19991 / 要約の誤りについて