arXiv論文メモ
新着一覧
cs.SD / cs.CL · 査読状況未確認

実際の会話から複数話者の声を抽出する難しさ

Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

Robert Sutherland, Stefan Goetze, Jon Barker

この論文をやさしく読む

ひとことで言うと

実会話では沈黙が多く登録音声も対象音声と異なるため、話者の声を取り出す手法の学習が難しくなることを調べた。

何に役立つ?

実際の複数人会話から話者の音声を抽出するモデルで、沈黙の偏りを考慮した損失関数の設計に役立つ。

この研究の面白いところ

沈黙が多い学習データ向けの損失関数で、STOIが0.55から0.60へ、別の音声品質指標が4.35から5.12へ改善した。

どこまで分かった?

要旨は登録音声との不一致も調べると述べるが、その影響の具体的な結果は記載していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

対象の話者や複数の話者の音声を抽出する技術は、他の話者や雑音が混ざる中から、望む話者の声を取り出す。ニューラルネットワークによる方法は、対象の発話と話者登録用の音声サンプルが量的に均衡し、しかもよく似ている人工データで学習・評価されることが多い。しかし実際の複数人の会話では、参加者は話している時間より黙っている時間のほうが長いことが多く、登録用サンプルと会話中の対象音声がかなり異なる場合もある。こうした要因は、実会話の録音による学習と評価に影響する。本研究は、学習時に沈黙が多すぎる影響を軽減する新しい損失関数を提案し、短時間客観了解度指標STOIを0.55から0.60へ、周波数重み付き区間信号対雑音比を4.35から5.12へ改善した。さらに、登録用音声と対象音声の不一致による影響を調べる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.

著者のコメント

Accepted to the International Workshop on Acoustic Signal Enhancement (IWAENC), Cremona, Italy, September 2026

arXiv ID: 2609.25948 / 要約の誤りについて