音声の主張を検索した証拠で検証するVeriSpeak
To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech
この論文をやさしく読む
ひとことで言うと
話された主張の真偽を音声言語モデルが確認できるか、検索した文章の証拠を使って評価する研究。
何に役立つ?
音声の事実検証システムを比較し、音声理解と証拠の照合を改善するための評価に役立つ。
この研究の面白いところ
3,879件の音声主張で、文章では正答するモデルも音声では失敗することを示す。検索と明示的な推論を組み合わせたモデルは86.1%の正解率に達した。
どこまで分かった?
正解率はベンチマークでの値であり、実際のニュースや政治発言全般での性能を示すものではない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
オンラインの誤情報は、ニュース映像、ポッドキャスト、インタビュー、政治演説、ソーシャルメディアの動画など、音声の形でも増えており、話された主張を直接検証する仕組みが必要になっている。本研究は、大規模音声言語モデル(LALM)による音声の事実検証を調べるベンチマークVeriSpeakを導入する。VeriSpeakには、時間、地理、関係性に関する3,879件の音声の主張が含まれ、真偽のラベルは均衡している。このベンチマークは、文章の事実検証能力が音声にも移るか、また検索を組み込んだLALMが文章の証拠を使って音声の主張を適切に裏付けたり反証したりできるかを調べるために設計された。実験では、一貫した文章と音声のモダリティ間の差が明らかになった。書かれた主張を安定して検証するLALMでも、同じ主張が音声になると失敗しやすい。さらに、検索した証拠と音声の主張をモデルがしばしば混同するため、検索だけによる改善は限られた。一方、検索と明示的な推論を組み合わせると主張と証拠の比較が改善し、思考向けに調整されたLALMは86.1%の正解率に達した。VeriSpeakは、音声の誤情報を有効に検出するには、音声の理解だけでなく、検索した証拠に基づく推論も必要であることを示す。データセットはHugging Faceの https://huggingface.co/datasets/abhiram4572/VeriSpeak で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Online misinformation increasingly appears in spoken formats such as news clips, podcasts, interviews, political speeches, and social media videos, creating a need for fact-checking systems that can verify claims directly from speech. We introduce VeriSpeak, a probe benchmark for studying speech-based fact verification in Large Audio Language Models (LALMs). VeriSpeak contains 3,879 spoken claims spanning temporal, geographical, and relational facts, with balanced true and false labels. The benchmark is designed to examine whether factual verification ability transfers from text to speech, and whether retrieval-augmented LALMs can use textual evidence to correctly support or refute spoken claims. Our experiments reveal a consistent text-speech modality gap: LALMs that verify written claims reliably often fail on the same claims when spoken. Moreover, retrieval alone provides limited gains because models frequently conflate retrieved evidence with the spoken claim. In contrast, retrieval combined with explicit reasoning improves claim-evidence comparison, with a thinking-tuned LALM reaching 86.1% accuracy. VeriSpeak highlights that effective speech misinformation detection requires not only speech understanding, but also grounded reasoning over retrieved evidence. The dataset is publicly available via Hugging Face at https://huggingface.co/datasets/abhiram4572/VeriSpeak.
著者のコメント
Accepted to EMNLP (Main) 2026
arXiv ID: 2609.30227 / 要約の誤りについて