文字起こしの曖昧さをトークン単位で許容する音声認識学習
A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition
この論文をやさしく読む
ひとことで言うと
音声だけでは一意に決まらない文字起こしの一部分を、単語全体ではなくトークン単位で学習時に許容する方法。
何に役立つ?
発音や表記の揺れを含む音声認識データで、正しい部分の教師信号を残しながら学習するのに役立つ。
この研究の面白いところ
19言語の25課題すべてでCTCを上回り、組み合わせた方式ではCTCより平均相対WERが9.45%低下した。曖昧な文字で回避確率が高まることも検証した。
どこまで分かった?
評価は三つのコーパスと記載された課題に基づく。その他の音声条件や運用環境への一般化は要旨からは判断できない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
自動音声認識の学習では通常、参照文字起こしだけを発話の正解ラベルとみなす。しかし逐語的とされる文字起こしでも、発音、つづり、語の表現に局所的な違いがあり、音響情報だけでは一意に決まらない場合がある。Omni-temporal Classification(OTC)は、コネクショニスト時系列分類(CTC)のアラインメントグラフにワイルドカード経路を加えてこうしたノイズを許容するが、単語単位の経路は粗すぎる。根拠の弱い一つのトークンを回避すると、単語全体に対する教師信号を捨ててしまうためである。 本研究はワイルドカード経路をトークン単位へ移し、根拠の弱いトークンだけを回避しながら単語の残りには教師信号を保つ。さらにトークン単位と単語単位の経路を、相補的な回避手段として組み合わせる。19言語、三つのコーパスにまたがる25課題すべてで、トークン単位のOTCはCTCを上回った。また、ワイルドカードの重みをエポック数で緩和する方式を、予測エントロピーに基づくスケジュールに置き換えた。この方式は同程度の性能を保ちながら学習期間への依存を減らす。スケジュールと両単位を組み合わせたグラフを併用すると、各コーパスで平均単語誤り率(WER)が最も低くなり、CTCに比べた平均相対WER低下率は9.45%だった。独立した検証者による文字起こしでは、トークン単位のモデルは、意見が分かれた文字でのワイルドカード回避確率をCTCより有意に高く割り当てた。これはトークン単位の許容が局所的な文字起こしの曖昧さを狙っていることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling, or lexical realization that the acoustics do not uniquely determine. Omni-temporal Classification (OTC) tolerates such noise by adding wildcard paths to the connectionist temporal classification (CTC) alignment graph, but its word-level arcs are too coarse, since bypassing one unsupported token discards supervision for the whole word. We move wildcard arcs to token granularity so unsupported tokens can be bypassed while the rest of the word stays supervised, and we combine token- and word-level arcs as complementary escape paths. Across 19 languages and three corpora, token-level OTC improves over CTC on all 25 tasks. We also replace epoch-indexed relaxation of the wildcard weights with a predictive-entropy-indexed schedule, which performs comparably while reducing dependence on training length. Combining this schedule with the hybrid graph gives the lowest mean word error rate (WER) on every corpus and a 9.45% average relative WER reduction over CTC. Independent validator transcriptions show that token-level models place significantly more wildcard-bypass probability than CTC on disputed characters, indicating that token-level tolerance targets localized transcript ambiguity.
著者のコメント
5 pages, 2 figures, 4 tables; submitted to ICASSP 2027
arXiv ID: 2609.30160 / 要約の誤りについて