arXiv論文メモ
新着一覧
eess.AS · 査読状況未確認

単語単位の強調を強化学習で制御する音声合成EmphTTS

EmphTTS: an emphasis-control TTS with reinforcement learning

Zirui Li, Rech Silas, Lauri Juvela, Tom Backstrom, Mikko Kurimo

この論文をやさしく読む

ひとことで言うと

合成音声で特定の単語を自然に強調できるよう、音の長さの予測器を強化学習で調整した。

何に役立つ?

読み上げ音声で話し手の意図を強調として伝える制御に役立つ可能性がある。

この研究の面白いところ

強調位置への報酬を音の長さの予測に与え、客観評価と人の選好評価の両方で改善を報告した。

どこまで分かった?

要旨には比較手法との差の具体的な数値や、多言語への適用結果は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

文章から音声を合成する技術では、文章入力に強調を明示的に指定しても、人間らしい強調を思いどおりに作ることは未解決であり、実際の用途で合成音声が意図を正確に伝えるうえで制約となる。近年、強化学習による音声合成システムの追加学習は人間の好みに合わせる方法として有望だが、従来法は単語単位の韻律制御に適用されていない。本研究は、単語単位の強調を直接最適化するため、強調位置の報酬を使うGroup Relative Policy Optimization(GRPO)を音の長さの予測器に適用した、非自己回帰型の音声合成システムEmphTTSを提示する。 評価では、EmphTTSが強調の制御可能性で最良となり、強調に関する客観評価でも最良の性能を示した。主観的な選好試験では、合成された正解音声と大半の比較手法より有意に好まれた。要素除去実験から、GRPOは教師あり微調整に基づく長さのモデル化や単純な話速調整を超えて強調の実現を改善し、別々に学習した長さ予測器と音声合成モデルの間の不一致も緩和することが分かった。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Generating controllable and human-like emphasis remains an open challenge in text-to-speech, even when explicit emphasis control signals are provided in the text input, limiting the communicative accuracy of synthetic speech in real-world applications. Reinforcement learning has recently shown promise for post-training TTS systems to align with human preference, yet existing methods have not been applied to word-level prosodic control. We present EmphTTS, a non-autoregressive TTS system that applies Group Relative Policy Optimization (GRPO) to the duration predictor with an emphasis localization reward, enabling direct optimization for word-level emphasis. Evaluations show that EmphTTS achieves the best emphasis controllability and performs the best in emphasis objective evaluation. In subjective preference tests, EmphTTS is significantly preferred over synthetic groundtruth and most baselines. Ablation studies show that GRPO improves emphasis realization beyond supervised-finetuning-based duration modeling and simple speaking-rate adjustment, while alleviating the mismatch between the independently trained duration predictor and TTS model.

著者のコメント

5 pages. Submitted to ICASSP 2027

arXiv ID: 2609.27599 / 要約の誤りについて