長い音声の専門用語を全体文脈で修正する方法
agentic-ger: terminology recovery in long-form speech using global context
この論文をやさしく読む
ひとことで言うと
長い音声の書き起こしで間違いやすい専門用語を、会話全体と元音声を使って修正する方法。
何に役立つ?
専門用語を含む長時間の音声認識結果を見直す用途が考えられる。要旨では中国語と英語の評価結果を報告している。
この研究の面白いところ
全文の文脈で候補を探し、必要な箇所だけ元音声を再認識して、確定した修正を次の判断にも使う。
どこまで分かった?
実験はGigaSpeechBench、4モデル、2認識システムで行われた。36.8%は中国語でWhisperとの比較における最大相対改善である。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声言語モデルの進歩により長い音声の自動認識は改善したが、専門分野の用語を正確かつ一貫して書き起こすことはなお難しい。著者らは、大規模言語モデルの知識と文脈を扱う能力を利用し、長い音声の用語を修正するエージェントAgentic-GERを提案する。エージェントは書き起こし全体の文脈から疑わしい用語を見つけ、曖昧な候補を判定する。候補の修正を確かめるため元の音声を選択的に再認識し、採用した修正をその後の判断に利用する。GigaSpeechBenchで4種類の大規模言語モデルと2種類の音声認識システムを用いて実験し、推論過程の有無にかかわらず、中国語と英語の両方で用語の改善が一貫して見られた。中国語音声ではWhisperを基準として、偏りを考慮した文字誤り率が最大36.8%相対的に低下した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent advances in speech language models have improved automatic speech recognition (ASR) for long-form audio. However, accurately and consistently transcribing domain-specific terminology remains challenging. Motivated by the world knowledge and contextual capability of large language models (LLMs), we propose Agentic-GER, an LLM-based agent for terminology correction in long-form speech. The agent uses global context from the full transcript to identify suspicious terms and resolve ambiguous hypotheses. It selectively re-transcribes the source speech to check candidate corrections, and uses accepted edits to guide subsequent decisions. Experiments with four LLMs and two ASR systems on GigaSpeechBench show consistent terminology improvements in both Chinese and English, with and without thinking. On Chinese speech, Agentic-GER achieves up to a 36.8% relative reduction in biased character error rate (B-CER) over the Whisper baseline.
著者のコメント
submitted to ICASSP 2027
arXiv ID: 2609.29428 / 要約の誤りについて