音声質問応答の誤答を見分ける不確実性指標を比較
Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs
この論文をやさしく読む
ひとことで言うと
音声への質問に答えるモデルの誤答を、不確実性の指標でどれほど見分けられるか比較した研究です。
何に役立つ?
音声質問応答システムの回答を点検する指標を選ぶ際、追加のモデル呼び出しが不要な方法と、その検出性能を比較できます。
この研究の面白いところ
選択肢式では単純な最初のトークンの確率が強く、音声を除くと誤答検出性能が質問文を除いた場合より大きく落ちました。
どこまで分かった?
結果は公開重みモデル4種類と指定した音声質問応答ベンチマークの評価です。自由記述式では平均正答率が下がり、検出指標の性能も選択肢式と同一ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声言語モデルは、音声に裏付けられない答えを自信ありげに出すことがある。このため、信頼できない回答を見分ける不確実性の推定が求められる。本研究では、公開重みのモデル4種類と音声質問応答ベンチマーク5種類について、確率に基づく指標、サンプリングに基づく指標、自己検証、証拠に基づく指標、対照的指標を比較した。選択肢式の評価では最初のトークンに基づく指標が全体として最も強く、最上位候補の確率は平均AUROC 0.740だった。10回のサンプリングによる離散的な意味エントロピーは0.708であり、前者はモデルへの追加呼び出しも必要としなかった。 4つのベンチマークで選択肢式から自由記述式に変えると、平均正答率は57.6%から36.6%に下がった。それでも不確実性は誤りの予測に有効で、意味エントロピー、最大トークンエントロピー、意味的一致度の平均AUROCは順に0.697、0.694、0.693だった。不確実性が回答に利用できる証拠を反映しているか調べるため、音声または質問文を除く入力アブレーションを行った。最上位候補の信頼度、エントロピー、サンプリングに基づく指標を通じて、音声を除くと誤答検出AUROCは平均0.101低下したのに対し、質問文を除いた場合の低下は0.010だった。これらの結果は、効率的な不確実性評価の基準を示し、音声言語モデルの不確実性が質問文よりも利用可能な音声の証拠に大きく依存することを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Audio-language models can produce confident answers unsupported by the audio, motivating uncertainty estimates that identify unreliable responses. We compare probability-based, sampling-based, self-verification, evidential, and contrastive measures across four open-weight models and five audio QA benchmarks. In multiple-choice evaluation, first-token measures are strongest overall, with top-1 probability achieving a mean AUROC of .740, compared with .708 for ten-sample discrete semantic entropy, while requiring no additional model calls. Across four benchmarks, shifting from multiple-choice to open-ended evaluation lowers mean accuracy from 57.6% to 36.6%, yet uncertainty remains predictive of errors: semantic entropy, maximum token entropy, and semantic agreement achieve mean AUROCs of .697, .694, and .693, respectively. To test whether uncertainty reflects the evidence available to answer the question, we perform input ablations that remove either the audio or the question. Across top-1 confidence, entropy, and sampling-based measures, removing audio reduces error-detection AUROC by .101 on average, compared with .010 when removing the question. Together, these results establish efficient uncertainty baselines and show that uncertainty in audio-language models depends substantially more on available audio evidence than on question text.
著者のコメント
Preprint
arXiv ID: 2609.28879 / 要約の誤りについて