ガーナの3言語で若者の健康相談向け音声認識を評価
Benchmarking and Domain Adaptation of Automatic Speech Recognition (ASR) for Adolescent Health Communication in Ghanaian Languages
この論文をやさしく読む
ひとことで言うと
ガーナの3言語で、若者の健康相談に使う音声認識を比較・適応し、音声アプリでも試した研究です。
何に役立つ?
考えられる用途は、対象言語での音声による健康情報へのアクセスです。アプリは地域の回答者50人で評価しました。
この研究の面白いところ
モデル比較、分野適応、実際のアプリでの検証をつなげ、エウェ語の単語誤り率を109.3%から64.8%に下げました。
どこまで分かった?
エウェ語でも単語誤り率64.8%と高く、要旨は対象分野の検証済みデータ不足を主な制約としています。医療上の正確さを保証した結果ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ガーナのトゥイ語、ダグバニ語、エウェ語による若者の健康相談を対象に、音声認識を端から端まで調べる。研究はつながった三段階で進む。第一に、各言語向けのWav2Vec2モデル三つと、複数形式を扱う言語モデルGemma 3n、Gemma 4の計五つを、一般分野の聖書音声資料と、若者の性と生殖に関する健康分野の音声認識データセットで比較し、文字誤り率と単語誤り率を測る。 第二に、比較結果に従い教師ありの分野適応を行う。追加訓練なしではGemma 4が最良だったが、その微調整は計算上実行できなかったため、小型のQwen3-ASR-0.6Bへ切り替えた。約9万標本の大きなガーナの聖書音声資料で微調整し、人が収集した対象分野の保留音声だけで評価した。微調整により全言語で単語誤り率が下がり、特にエウェ語では109.3%から64.8%へ44.5ポイント下がり、文字誤り率も65.1%から24.9%へ下がった。 第三に、3言語で実際に稼働する音声中心の健康相談アプリKasaHealthと、技術評価用のSenti-Checkを通じて検証した。KasaHealthを地域の回答者50人で試したところ、対話の承認率は100%、翻訳を「良い」または「非常に良い」とした割合は72%、推薦したいと答えた割合は92%だった。同時に、実使用を最も制約する分野固有のデータ不足が浮き彫りになった。三段階の証拠は、これらの言語で主な制約となるのがモデルの能力や計算量ではなく、検証済みの対象分野データであることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
This paper presents an end-to-end study of automatic speech recognition (ASR) for adolescent health communication in three Ghanaian languages (Twi, Dagbani, and Ewe). The work proceeds in three connected stages; First, we benchmark five ASR systems (three language-specific Wav2Vec2 models and two multimodal LLMs, Gemma 3n and Gemma 4) on a general-domain Bible corpus and a Youth Adolescent Sexual and Reproductive Health (ASRH) Domain ASR dataset, using Character and Word Error Rate (CER, WER). Second, guided by the benchmark, we perform supervised domain adaptation: although Gemma 4 was the strongest zero-shot candidate, fine-tuning it proved computationally infeasible, so we pivoted to the compact Qwen3-ASR-0.6B, fine-tuned on a large Ghana Bible corpus (~90k samples) and evaluated strictly on held-out human-collected in-domain audio. Fine-tuning reduced WER on every language, most dramatically for Ewe (WER from 109.3% to 64.8%, a drop of 44.5 pp; CER from 65.1% to 24.9%). Third, we validate the work through KasaHealth, a live voice-first ASRH application deployed in all three languages, complemented by Senti-Check, a technical evaluation harness. KasaHealth was tested by 50 community respondents and achieved a 100% chat-approval rate, a 72% Good-or-Excellent translation rating, and a 92% would-recommend rate, while surfacing the domain gaps that most constrain real-world use. Across all three stages the evidence converges: for these languages the binding constraint is validated in-domain data, not model capability or computation.
著者のコメント
34pages, 8figures,
arXiv ID: 2609.29798 / 要約の誤りについて