arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

147言語・110万時間超の高音質音声データ

YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

William Chen, Shinnosuke Takamichi, Sayaka Shiota, Satoru Fukayama, Samuele Cornell, Shinji Watanabe

この論文をやさしく読む

ひとことで言うと

147言語で110万時間超の高音質な音声を集めた、音声研究用の公開データセット。

何に役立つ?

多言語音声認識や音声符号化モデルの学習と評価に使える。

この研究の面白いところ

48 kHzの多チャンネル音声を大規模に収集し、言語の分布だけでなく音声と文字起こしの品質も分析した。

どこまで分かった?

ラベルは弱いラベルであり、要旨では文字起こし品質の具体的な値は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

著者らは、147言語の48 kHz多チャンネル音声を110万時間超含む、弱いラベル付きの音声コーパスYODAS v3を提示する。CC BY 3.0ライセンスで公開される。著者らによれば、これまでで最大の公開音声データセットであり、高忠実度のステレオ音声を備えた初の真に大規模な音声コーパスでもある。まずコーパスの収集方法を示し、言語間の量の偏りを抑えて音声データを集める新しい方法を導入する。その効果は収集データの言語分布に表れ、22言語で1万時間超、73言語で5千時間超のデータがある。次に言語分布、音声の品質、文字起こしの品質など、データの構成を広く分析する。最後に、基準となる音声認識モデルとニューラル音声符号器を学習し、データセットの有効性を示す。データは公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

We present YODAS v3, a weakly-labeled speech corpus containing over 1.1 million hours of 48kHz multi-channel audio in 147 languages, released under a CC BY 3.0 license. YODAS v3 is not only the largest open speech dataset to date, but also the first truly large-scale speech corpus with high-fidelity stereo audio. We first provide the collection methodology for the corpus, where we introduce new techniques for gathering language-balanced speech data. The effectiveness of our approach is shown by the language distribution of the crawled data: 22 languages in YODAS v3 have over 10K hours and 73 languages have over 5K hours of data. We then conduct extensive analyses on the composition of the data, such as the distribution of languages, audio quality, and transcription quality. Finally, we train baseline speech recognition and neural codec models to show the effectiveness of the dataset. Download at https://huggingface.co/datasets/espnet/yodas3.

著者のコメント

Interspeech 2026; 6 Pages

arXiv ID: 2609.29448 / 要約の誤りについて