arXiv論文メモ
新着一覧
eess.AS / cs.AI / cs.CL / cs.SD · 査読状況未確認

農業相談の音声認識をモデル交換なしで改善する処理手法

Model-Agnostic and Language-Agnostic Voice Pipeline Improvement for the Agriculture Domain

Aakash Singh, Lakshmi Pedapudi, Chandrashekar M S, Sanyam Singh, Naga Ganesh, Vineet Singh

この論文をやさしく読む

ひとことで言うと

農家の音声相談を、雑音除去・話者選択・農業語彙での訂正を組み合わせて書き起こしやすくします。

何に役立つ?

機械音や他人の声が混じる現場録音で、作物・害虫・薬剤・数量などの重要語を正しく扱うための方法です。

この研究の面白いところ

既存の音声認識モデルを置き換えず前後処理を追加します。三言語の評価で、複数話者がいる録音ほど話者選択の効果が大きくなりました。

どこまで分かった?

評価はHindi・Telugu・OdiaのFarmerChat録音です。クラウドモデルのWER16〜23%低減は相対値で、調整を行ったのは話者分離の段階です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

FarmerChatは、Digital Greenが小規模農家に提供するAI農業助言アシスタントであり、利用者は自分の言語のテキスト、音声、写真を通じて利用する。音声はこの利用者層にとって重要な経路だが、現場で録音された発話は汎用の自動音声認識(ASR)にとって難しい。録音には、機械の騒音、背景で流れるメディアの音、他の話者の発話、農業特有の語彙が頻繁に含まれるからである。こうした条件は、農家の質問の意味を担う作物、害虫、薬剤、数量の語に、とりわけ大きな影響を及ぼす。 本研究では、基礎となるASRモデルを微調整したり置き換えたりせずに、FarmerChatの音声認識品質を改善する、モジュール型でモデルに依存しないパイプラインを提案する。このパイプラインは、適用を判定するゲート付き音声強調、話者区間分割と対象話者の選択、ASR、重み付き農業用語辞書による分野を考慮した訂正、信頼できない文字起こしを検出する品質ゲートを組み合わせる。微調整するのは話者区間分割の段階のみで、それ以外は共通のインターフェースを介して既成モデルを利用する。 人手で注釈を付けたヒンディー語、テルグ語、オディア語のFarmerChat録音を使い、単語誤り率(WER)と、農業用語をより重視する分野重み付き誤り率で評価する。最大の改善は複数話者の録音で見られ、対象話者の選択により他者の発話が文字起こしへ混入することを防げる。全コーパスでは、3つのクラウドASRモデルでWERが相対的に16〜23%、端末内モデルで5%低下した。複数話者の録音では、低下率はクラウドモデルで32〜42%、端末内モデルで16%だった。報告した低下はいずれも統計的に有意である。これらの結果は、基礎となるASRモデルを維持しつつ、対象を絞った前処理、話者選択、分野を考慮した後処理によって、農業音声の文字起こしを大幅に改善できることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

FarmerChat is Digital Green's AI-powered agricultural advisory assistant for smallholder farmers, who access it in their own language through text, voice, or photographs. Voice is a critical channel for this population, yet field-recorded speech is challenging for general-purpose automatic speech recognition (ASR) because recordings frequently contain machinery noise, background media, competing speakers, and domain-specific agricultural vocabulary. These conditions disproportionately affect crop, pest, chemical, and quantity terms that carry the meaning of a farmer's query. We present a modular, model-agnostic pipeline for improving ASR quality in FarmerChat without fine-tuning or replacing the underlying ASR model. The pipeline combines gated audio enhancement, speaker diarization and target-speaker selection, ASR, domain-aware correction using a weighted agricultural lexicon, and a quality gate for detecting unreliable transcripts. Only the diarization stage is fine-tuned; all other stages use off-the-shelf models behind common interfaces. We evaluate the pipeline on human-annotated FarmerChat recordings in Hindi, Telugu, and Odia using word error rate (WER) and a domain-weighted error rate that gives greater importance to agricultural terminology. The largest improvements occur on multi-speaker recordings, where target-speaker selection prevents competing speech from entering the transcript. Across the full corpus, the pipeline reduces WER by 16-23% relative on three cloud ASR models and by 5% on an on-device model. On multi-speaker recordings, the reductions are 32-42% for the cloud models and 16% for the on-device model. All reported reductions are statistically significant. These results show that targeted preprocessing, speaker selection, and domain-aware post-processing can substantially improve agricultural speech transcription while preserving the underlying ASR model.

著者のコメント

20 tables, 11 figures, 23 pages

arXiv ID: 2609.20504 / 要約の誤りについて