方言と英語混じりの発話を含むベトナム語音声・偽音声データ集
VietPrism: A large-scale Vietnamese speech and deepfake corpus with diverse dialects and code-switching
この論文をやさしく読む
ひとことで言うと
話者、方言、英語混じりの発話を含むベトナム語音声と、同じ話者・文章に対応させた偽音声をまとめたデータ集を作った研究。
何に役立つ?
ベトナム語の音声認識や偽音声検出を、方言や話者の条件を分けて評価する際に役立つ。提供されたデータを使えば、同じ内容・話者の実音声と偽音声も比較できる。
この研究の面白いところ
実際の発話993.4時間に加えて偽音声を3100時間以上作り、語彙と話者をそろえた対を設けている。検出器の誤り率が話者の類似度や方言で変わる点も示した。
どこまで分かった?
要旨の検出性能は学習済み多言語検出器五つのゼロショット評価による。すべての検出器や生成方法で同じ誤り率になるとは述べていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ベトナム語音声の研究では、音声認識と、話者、方言、言語の切り替え、ディープフェイクの解析を別々に扱うデータ資源が制約となっている。本研究は、これらの要素を大規模にまとめた、公開の複数領域にまたがるコーパスVietPrismを導入する。実世界の動画8388本から確認済みの話者1262人について、実際の発話40万3941件、計993.4時間を収録する。著者らの知る限り、文字起こし、一貫した話者識別子、五つの方言群、自然に生じたベトナム語と英語の切り替えを併せて提供する初めての大規模ベトナム語コーパスである。英語との言語切り替えは収録時間のほぼ半分を占める。 さらに、オープンソースと商用の四つの音声合成システムで、3100時間を超える偽音声を作成する。それぞれの偽音声は確認済み話者の参照音声を条件として生成し、文字起こしと話者が一致する実際の発話と対にしている。これにより、語彙や話者の違いによる交絡を減らした統制された評価が可能になる。学習済み多言語検出器五つをゼロショットで評価すると、頑健性の不足が明らかになった。等誤り率(EER)は検出器と生成器の組み合わせによって大きく変わり、最近の多言語検出器DFA-1Bでは話者の類似度が増すにつれ16.3%から33.6%へ悪化した。方言別の結果でも、モデルによって異なる格差が見られる。自然な言語的多様性と統制された偽音声生成を一つにすることで、VietPrismはベトナム語音声モデルと信頼できる音声ディープフェイク検出を研究するための難しい基盤を提供する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Vietnamese speech research is constrained by resources that isolate automatic speech recognition from speaker, dialect, code-switching, and deepfake analysis. We introduce VietPrism, an open, multi-domain corpus that brings these dimensions together at scale: 993.4 hours and 403,941 bona fide utterances from 1,262 verified speakers across 8,388 real-world videos. To our knowledge, it is the first large-scale Vietnamese corpus to jointly provide transcripts, consistent speaker identities, five dialect groups, and naturally occurring Vietnamese--English code-switching, which constitutes nearly half of the corpus by duration. We further create over 3.1K hours of spoof speech with four open-source and commercial synthesis systems. Every spoof is conditioned on a verified speaker reference and paired with a transcript- and speaker-matched bona fide utterance, enabling unique controlled evaluation with reduced lexical and identity confounds. Zero-shot evaluation of five pretrained multilingual detectors reveals striking brittleness: EER greatly varies across detector--generator pairings, while recent multilingual detector DFA-1B degrades from 16.3% to 33.6% as speaker similarity increases. Dialect-stratified results expose further model-dependent disparities. By unifying natural linguistic diversity with controlled spoof generation, VietPrism provides a challenging foundation for Vietnamese speech modeling and trustworthy audio-deepfake detection.
著者のコメント
Preprint for ICASSP 2027 submission
arXiv ID: 2609.30005 / 要約の誤りについて