16言語の長時間音声を理解するモデルの評価
MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing
この論文をやさしく読む
ひとことで言うと
長い録音の理解を、言語や音の種類、必要な作業ごとに分けて評価するベンチマークを作った。
何に役立つ?
多言語の音声モデルが、どの条件で聞き取りや時刻合わせを誤るか調べるのに役立つ。
この研究の面白いところ
16言語、8領域、計1,377.9時間の実録音を使い、音声外の手掛かりも含めて失敗を分析した。
どこまで分かった?
長距離の情報検索は比較的強い一方、正確な時刻合わせや自然な音の事実への根拠付けには弱さが残る。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
長時間音声の性能を文脈長や全体の正答率だけでまとめると、言語、証拠、課題が難しさにどう関わるかが隠れる。MuLA-Benchは、実環境で録音された1,769件、計1,377.9時間の音声に対する自由回答の5,038問で、16言語と8領域を含む。言語と領域を均衡させた意味理解の課題は条件をそろえた比較を可能にし、補完的な音響課題には自然に生じた非音声の証拠を残す。証拠に基づく設問生成、安易な手掛かりへの依存の検査、言語の専門家による確認により、共通の原音源を翻訳したり、目標の音を人為的に加えたりせず、検証可能な設問を作る。十の音声・言語モデルを評価し、固定した八モデルの集団について診断を行った。言語の順位は領域や課題で変わり、音響と意味理解の成績差は求める処理によって異なり、正しい出来事を見つけても時間的な誤りが残ることがある。長い範囲からの情報検索は比較的良好だが、正確な時刻合わせと自然な音響事象についての事実に基づく回答は弱い。単一の長文脈スコアでは捉えられない条件付きの失敗パターンを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape difficulty. We introduce MuLA-Bench: 5,038 open-ended questions over 1,769 in-the-wild recordings totaling 1,377.9 hours, covering 16 languages and eight domains. A balanced Language x Domain semantic track supports controlled comparisons, while a complementary acoustic track preserves naturally occurring non-speech evidence. Evidence-grounded generation, shortcut checks, and language-expert review provide auditable questions without translating a shared source set or injecting target sounds. We evaluate ten audio-language models and conduct pooled diagnostics on a fixed eight-model cohort. Language rankings change across domains and tasks; acoustic-semantic performance gaps vary with the requested operation; and temporal errors can persist after the correct event is identified. Long-range retrieval is comparatively strong, while precise clock alignment and factual grounding of natural acoustic events remain fragile. MuLA-Bench thus exposes conditional failure patterns that a single long-context score does not capture.
arXiv ID: 2609.23416 / 要約の誤りについて