音声認識モデルの圧縮が話者間の修正負担に与える差
Temporal Taxation Compounds Under Post-Training Compression of Whisper Models
この論文をやさしく読む
ひとことで言うと
音声認識モデルを軽くすると、話者集団によって文字起こしの誤りを直す手間がどう変わるかを調べた研究。
何に役立つ?
端末向け音声認識を圧縮して配備する前に、集団別の誤り率と利用者の修正時間を評価する必要性を示す。
この研究の面白いところ
50%の枝刈りで集団間の推定修正時間差が音声1分当たり30秒から64秒になった一方、蒸留は27条件中21条件で差を縮めた。
どこまで分かった?
修正時間への換算は誤り1件につき5秒という仮定に基づく。結果は調べたWhisper系列と三つのデータセット、圧縮条件についてのもの。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声認識モデルの属性別の公平性は通常、完全精度の状態で監査される。しかし実際に配備されるモデルは量子化、枝刈り、蒸留を経ている。本研究は、音声やその特徴表現ではなくモデルの重みを変える学習後の圧縮によって、属性集団ごとの誤りの負担が変わるかを問う。Fair-Speech、Common Voice 25、AfriSpeech-200でWhisper系列を調べた。Whisper-large-v3に50%のWanda枝刈りを施すと、Fair-Speechでは黒人・アフリカ系アメリカ人(Black/AA)とアジア人の間の時間的負担の差が大きく広がった。最もサービスの悪い集団と良い集団の単語誤り率の絶対差は2倍以上になった。転記誤り1件の修正に5秒かかると仮定すれば、音声1分当たりの修正時間の差は30秒から64秒に増える。この相対増加率111%は、1件当たりの修正費用の仮定に左右されず、音質を調整しても残った。ビームサーチによる復号では一部軽減されるが、なお86%の増加が残る。 エッジ端末向けのモデルサイズでは、INT4 HQQ量子化によって西アフリカのアクセントで転記文が破綻して繰り返す現象が5~7倍に増えた。一方、蒸留は評価した27条件(教師・生徒の組、精度、データセット)のうち21条件で属性間の差を縮め、例外は一つのモデルの組に集中した。ChoiとChoi(2025)の「時間的負担」の概念を数量指標として定式化し、完全精度モデルを一時点だけで監査しても、圧縮後の配備時に、既に不利な話者へ生じる負担を捉えられないことを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Automatic speech recognition models are audited for demographic fairness at full precision, yet the models that ship to production have been quantized, pruned, and distilled. We ask whether post-training weight compression, which alters model weights rather than the audio signal or its feature representation, redistributes error burden across demographic groups. Across the Whisper family on Fair-Speech, Common Voice 25, and AfriSpeech-200, 50% Wanda pruning of Whisper-large-v3 sharply widens the Black/AA-vs-Asian temporal-taxation differential on Fair-Speech: the absolute word-error-rate gap between the worst- and best-served groups more than doubles; at an assumed cost of five seconds of correction effort per transcription error this is a rise from 30 to 64 seconds of correction time per minute of speech. This +111% relative increase is invariant to the assumed per-error cost, survives an audio-quality control, and is only partly mitigated by beam-search decoding, which still leaves an +86% increase. At edge model size, INT4 HQQ quantization compounds catastrophic transcript loops on West African accents by factors of five to seven. Distillation, by contrast, narrows demographic gaps in 21 of 27 evaluated settings (teacher-student pair, precision, and dataset), with the exceptions concentrated on a single model pair. We cast the temporal-taxation construct of Choi and Choi (2025) as a quantitative metric, and show that single-snapshot fairness audits on full-precision models do not capture the deployment-time burden that compression places on already-marginalized speakers.
著者のコメント
Accepted to IMPACT-SPEECH @ EMNLP 2026
arXiv ID: 2609.28739 / 要約の誤りについて