発話の実現方法を指定できるフローマッチング音声合成
ReaFlow-TTS: Realization-Conditioned Flow Matching for High-Quality and Controllable Speech Synthesis
この論文をやさしく読む
ひとことで言うと
文章から音声を作る際に、発話ごとの違いを潜在変数で表し、感情に関わる属性を段階的に調整する手法。
何に役立つ?
考えられる用途は、目標音声を毎回用意せずに話し方を制御する音声合成である。要旨では合成品質と属性操作を評価している。
この研究の面白いところ
一つの潜在変数を生成過程全体で使い、初期ノイズが違っても音高や長さなどの傾向が再現される。
どこまで分かった?
要旨には基準手法との実験と主観評価が述べられるが、言語や話者を広く変えた場合の適用範囲は記されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
フローマッチングによるテキスト音声合成では、同じ生成条件でも、異なる発話の実現形によって目標とする速度が変わり得る。二乗誤差で学習した決定論的な速度場は、その条件付き平均を予測するため、実現形に依存する変動を平均化してしまう。一方、こうした変動をモデル化するだけでは、意味を解釈できる形で属性を操作する手段が自動的に得られるわけではない。本研究はReaFlow-TTSを提案する。これは、発話全体に対応する確率的な実現形の潜在変数を導入し、生成の全過程を通じて速度予測をその変数で条件付けるフローマッチングの枠組みである。さらに、実現形の空間に快・不快、覚醒度、支配性(VAD)の意味付けを課すことで、推論時に目標音声を与えなくても属性を直接、段階的に操作できるようにする。 実験では、条件をそろえた全面マスクの基準手法より音声合成の品質が改善した。また、初期ノイズの標本を変えても、潜在変数による音高、エネルギー、時間的な傾向が再現され、潜在変数が再利用可能な発話実現条件として使われていることを振る舞いの面から示した。主観評価でも、生成条件を変えた場合にVAD属性を段階的に操作でき、自然さの変化は比較的小さいことが示された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
In flow-matching text-to-speech (TTS), different speech realizations can induce different target velocities under the same generation conditions. A deterministic velocity field trained with squared error predicts their conditional mean, thereby marginalizing realization-dependent variation. Meanwhile, modeling such variation does not inherently provide a semantically interpretable interface for attribute manipulation. We propose ReaFlow-TTS, a realization-conditioned flow-matching framework that introduces an utterance-level stochastic realization latent and uses it to condition velocity prediction throughout the generation trajectory. We further impose valence-arousal-dominance (VAD) semantics on the realization space, enabling direct and graded attribute manipulation without target speech at inference. Experiments demonstrate improved synthesis quality over a matched full-mask baseline and reproducible latent-induced pitch, energy, and timing tendencies across initial-noise samples, providing behavioral evidence that the latent is used as a reusable realization condition. Subjective evaluation further demonstrates graded VAD manipulation across generation contexts with only modest changes in naturalness.
arXiv ID: 2609.28906 / 要約の誤りについて