arXiv論文メモ
新着一覧
cs.CV / cs.AI · 査読状況未確認

動画エージェントの誤答が生じる段階を介入実験で特定

Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models

Shuzhi Gong, Fengze Sun, Yuansan Liu

この論文をやさしく読む

ひとことで言うと

動画エージェントの誤答が、時間的な位置特定、観察、推論のどこから広がるかを調べています。

何に役立つ?

動画エージェントの評価で、段階ごとの得点が実際の回答の信頼性を表すか見直す際に役立ちます。要旨では三つの構成での介入実験を報告しています。

この研究の面白いところ

下流の課題を固定しながら途中の出力だけを置き換え、位置特定の誤りが観察の誤りより大きく効くことを因果的に比較しています。

どこまで分かった?

要旨にある比較は三つのエージェント構成での60,008回の実行に基づきます。すべての動画課題や構成で同じ比率になるとは述べていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

動画の理解には、時間的な位置特定、視覚的な観察、推論を分けて処理する多段階の大規模言語モデルエージェントが使われることが増えている。しかし、各段階は通常、異なるベンチマークやデータ分布で評価されるため、事実と異なる回答がどこで生じるか判断しにくい。本研究はまず、既存のベンチマークをこれらの段階ごとに整理し、得点が診断指標として一貫しないことを示す。ある段階での性能が高くても下流での誤答が少ないとは限らず、同じ能力を対象とするベンチマーク同士でも評価が食い違う。 そこで、下流の課題を固定したまま個々の段階の出力を置き換える、因果的な段階介入の手順を導入する。三つの動画エージェント構成で計60,008回実行した結果、時間的な位置特定が下流の誤りの主な原因で、その出力を壊したときの因果的な影響は視覚観察を壊した場合のおよそ4倍だった。位置特定の成功には、時間区間の厳密な重なりより正しい領域を見つけることが重要であり、標準的なmIoU指標が下流の信頼性を予測しにくい理由になる。また、証拠が欠けている場合より、誤った証拠がある場合のほうが大幅に有害だった。既存ベンチマークをこの介入結果と照合すると、その得点は連鎖的な誤りへの因果的な感度を確実には予測できず、分布が変わると機能しない場合もあった。これらの結果は、信頼できる動画エージェントには、段階ごとの介入に基づく評価が必要であることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing benchmarks around these stages and show that their scores provide inconsistent diagnostic signals: stronger stage-level performance does not reliably imply lower downstream hallucination, and even benchmarks targeting the same capability can disagree. We therefore introduce a causal stage-intervention protocol that overwrites individual stages while holding the downstream task fixed. Across 60,008 runs on three video-agent architectures, we find that grounding is the dominant source of downstream error, with roughly four times the causal impact of corrupting visual observations. Successful grounding depends primarily on locating the correct region rather than precise temporal overlap, explaining why standard mIoU metrics poorly predict downstream reliability. We further find that incorrect evidence is substantially more harmful than missing evidence. Finally, auditing existing benchmarks against these interventions reveals that their scores do not reliably predict causal cascade sensitivity and can fail under distribution shift. These results motivate intervention-based, stage-aware evaluation for trustworthy video agents.

著者のコメント

Accepted in NeurIPS 2026 TAE workshop

arXiv ID: 2609.28991 / 要約の誤りについて