arXiv論文メモ
新着一覧
cs.CL / eess.AS · 査読状況未確認

テルグ語の音声質問応答と自動評価の信頼性を検証

VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering

Bhavana Akkiraju, Ravi Sastry Kolluru, Sri Charan D, Srihari Bandarupalli, Santosh Kesiraju, Anil Vuppala

この論文をやさしく読む

ひとことで言うと

テルグ語の音声質問にモデルが答えられるか、また自動採点が人の判断に合うかを調べる評価データセットです。

何に役立つ?

低資源言語の音声QAで、音声認識、翻訳、回答生成、採点のどこに問題があるかを検討する材料になります。

この研究の面白いところ

六領域の2,001問と2.53時間の音声、人が確認した答えを用意しています。参照と表現が違う正答を不当に減点する採点器の傾向も確認しています。

どこまで分かった?

Geminiによる採点が人に最も近かったものの、厳しさは均一ではありません。音声の音韻的混同や翻訳による文化的情報の消失も観察され、自動評価を無条件に正解扱いできない結果です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルによって質問応答は急速に進歩してきたが、テキスト・音声のいずれにおいても、その中心は高資源言語である。テルグ語の音声質問応答(SQA)のベンチマークは未開拓であり、この設定での自動評価の信頼性も定量化されていない。本研究では、6分野にわたる2,001組の事実を問う質問・回答、2.53時間の音声、2言語の書き起こし、人手で検証した参照回答からなるテルグ語SQAベンチマークVākQAを導入する。 まず、人間の判断に照らして評価手法を検証する。Geminiを判定器とする方法は人間の評価に最も近いが、厳しさは一様ではない。一方、重み公開の判定モデルは、参照回答と表現が異なる正しいテルグ語回答を系統的に低く評価する。この検証済みの設定で、入力モダリティ、言語、分野を変え、商用モデルと重み公開モデルを比較評価する。 テルグ語の言い回しは翻訳で失われる文化的な固有性を保持すること、音声入力は質問の意味を変える音韻的な混同を招くこと、音声認識(ASR)と機械翻訳(MT)を直列につなぐと誤りが段階的に積み重なることを観察する。VākQAは一般公開する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.

著者のコメント

Paper is accepted in IEEE SLT 2026

arXiv ID: 2609.19879 / 要約の誤りについて