長さを事前指定しないゼロショット音声生成・編集EditVoice
EditVoice: Variable-Length Non-Autoregressive Zero-Shot TTS and Speech Editing with Edit Flows
この論文をやさしく読む
ひとことで言うと
生成する音声の長さを先に決めず、音声の挿入や削除も含めて並列に合成・編集するモデルです。
何に役立つ?
例示音声に合わせた音声合成や、テキストに基づく音声の編集に使えると考えられます。要旨では音声合成と編集のベンチマークで評価しています。
この研究の面白いところ
前方または後方の音声を手がかりに欠けた部分を補い、生成した音声も追加学習なしで再編集できます。
どこまで分かった?
競争力があるという定性的な報告で、要旨には具体的な指標値や各比較手法との差は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
近年の非自己回帰型のゼロショット音声合成モデルは並列に音声を生成できるが、通常は生成前に出力列の長さを指定しなければならない。本研究は、著者らの知る限り初の可変長の非自己回帰型ゼロショット音声合成モデルEditVoiceを導入する。Edit Flowsを使い、挿入、削除、置換を通じて音声内容と系列長を同時に更新する。EditVoiceは、音声の欠損部分を補う訓練を採用し、ゼロショットの音声合成とテキストに基づく音声編集を統一する。推論時には音声の手がかりを前方にも後方にも置ける。さらに、二つの配置から生じる相補的なEdit Flowの予測を利用するComplementary Prompt Sampling(CPS)を導入する。 EditVoiceは、訓練時の音源を超えて、元の音声やモデルが生成した音声も編集できることが分かった。この一般化を、一連の処理を通した編集と、生成後に追加学習なしで行う改良に利用する。GigaSpeechの1万時間のデータで訓練したEdit Flowモデルは、Seed-TTS Eval ENとLibriSpeech-PCで競争力のあるゼロショット音声合成性能を、RealEditで音声編集性能を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent non-autoregressive (NAR) zero-shot text-to-speech (TTS) models generate in parallel but typically require the target sequence length to be specified before generation. We introduce EditVoice, to our knowledge the first variable-length NAR zero-shot TTS model, which uses Edit Flows to jointly update speech content and sequence length through insertions, deletions, and substitutions. EditVoice adopts speech-infilling training, which unifies zero-shot TTS and text-based speech editing and allows both prefix and suffix speech prompt placements at inference. We introduce Complementary Prompt Sampling (CPS) to leverage the complementary Edit Flow predictions induced by the two prompt placements. We further find that EditVoice can edit source and model-generated speech beyond its training sources. We use this generalization for end-to-end editing and training-free post-generation refinement. With the Edit Flow model trained on 10K h of GigaSpeech, EditVoice demonstrates competitive zero-shot TTS performance on Seed-TTS Eval EN and LibriSpeech-PC and speech editing performance on RealEdit.
著者のコメント
5 pages, 3 figures. Audio samples: https://dhy02.github.io/editvoice-demo/
arXiv ID: 2609.29889 / 要約の誤りについて