音素の似た発話の集まりから音声合成の学習データを選ぶ
Structure Before Sampling: Community-Aware Core-Set Selection for Data-Efficient Text-to-Speech
この論文をやさしく読む
ひとことで言うと
音声合成の学習に使う発話を、似た音素配列のグループから偏りなく選び、少ないデータでも音の組み合わせを広く覆います。
何に役立つ?
音声時間や学習時間の予算が限られる場合に、学習用の発話を選ぶ方法として役立ちます。既存コーパスでの選択結果が中心です。
この研究の面白いところ
発話のグラフに強いコミュニティ構造があることを先に確かめてから、その構造をデータ選択に使っています。希少な音素の組み合わせにも着目します。
どこまで分かった?
評価言語はベンガル語と英語です。全コーパスを上回る3.93%対4.47%の比較はベンガル語で、同じエポック数という条件付きです。全言語や同じ総更新回数での優位性ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声合成(TTS)のコーパスは収録費用が高い一方、多くの発話は新しい音声学的情報をほとんど追加しない。コアセット選択は、音声の総時間を固定した予算の下で小さな学習部分集合を選び、この費用を減らす。本研究では、各発話を音素的に最も似た発話に結び付ける音素配列グラフとしてコーパスを表し、まずこのグラフに構造があるかを調べる。ベンガル語と英語のコーパスで、そのクラスタリングは同規模のランダムグラフのそれぞれ199倍と56倍となり、モジュラリティは次数を保つランダムグラフの2倍を超えた。 次に、Community Representativeという選択法を提案する。希少な音素を多く含む発話から始め、グラフのコミュニティをまたいでサンプリングし、各コミュニティ内でも選択を分散させる。両言語のすべての予算で、ランダム選択やエントロピーに基づく選択より多くの希少な音素2連続を覆い、この優位性は取り置いた発話でも維持された。選んだ20%のコアセットで学習したTTSモデルは、同じ音声時間のランダム部分集合やエントロピーに基づく部分集合で学習したモデルより、両言語で有意に低い文字誤り率(CER)を示した。すべてのモデルを同じエポック数で学習した場合、ベンガル語のコアセットモデルは全コーパス学習も上回り、CERは3.93%対4.47%となり、学習時間は4.5分の1だった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Text-to-speech (TTS) corpora are costly to record, yet many utterances add little new phonetic information. Core-set selection reduces this cost by choosing a small training subset under a fixed audio-duration budget. We represent a corpus as a phonotactic graph that links each utterance to its most phonemically similar ones, and we first test whether this graph has structure. In Bangla and English corpora, its clustering is 199 and 56 times that of a size-matched random graph, and its modularity is more than twice that of a degree-preserving random graph. We then propose Community Representative, a selector that samples across graph communities and spreads its choices within each one, starting from utterances rich in rare phonemes. At every budget and in both languages, it covers more rare phoneme bigrams than random and entropy-based selection, and this lead holds on held-out utterances. TTS models trained on its 20% core-sets have a significantly lower character error rate (CER) than models trained on equal-duration random or entropy-based subsets in both languages. When all models train for the same number of epochs, the Bangla core-set model also outperforms full-corpus training (3.93% vs. 4.47% CER) with 4.5x less training time.
arXiv ID: 2609.24275 / 要約の誤りについて