会話に応じて話し方を変える音声合成手法
Interactive TTS: Dynamic Speaking Style Adaptation for Expressive Speech Synthesis
この論文をやさしく読む
ひとことで言うと
対話の状況に合う話し方を選びつつ、話者の声質を保つ音声合成手法を提案した。
何に役立つ?
対話型の読み上げで、会話の流れに合わせた話し方を生成する用途が考えられる。実際の利用環境での効果は要旨からは分からない。
この研究の面白いところ
話し方の決定を実行可能な指示として明示し、音声生成との間を2種類の学習方法でつないでいる。
どこまで分かった?
VStyleとSpeechParaling-Benchで従来モデルを上回ったという実験結果が示される。要旨には改善幅や実環境での評価は記載されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
複数ターンのマルチモーダルな対話に合わせて話し方を動的に変えることは、音声合成(TTS)における大きな課題である。既存の文脈考慮型TTSは通常、対話の文脈から音声へ一貫して写像する。この暗黙的なモデル化では話し方の判断を教師あり学習で扱いにくく、話し方・声質・発話内容が絡み合うため、指示への追従が弱くなり、ターンをまたいで声質が大きくずれることがある。本稿は、文脈に合い、話者の声質を保った音声を生成する、話し方を動的に調整する枠組みInteractive TTSを提案する。文脈に基づく話し方の判断を、実行可能な指示として明示的にモデル化して処理を分離する。話し方の判断と音声生成をつなぐため、反復的な棄却サンプリングによる微調整(Iterative RSFT)と、文脈考慮型の直接選好最適化(CADPO)を導入し、指示追従と会話文脈への音声の適合を大きく改善した。広範な実験で、VStyleとSpeechParaling-Benchにおいて従来の最先端モデルを上回った。デモは https://wjtian-wonderful.github.io/InteractiveTTS/ で公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Dynamic speaking style adaptation in multi-turn multimodal interaction remains a major challenge for text-to-speech (TTS) systems. Existing context-aware TTS (CTTS) methods typically map dialogue context to speech in an end-to-end manner. Such implicit modeling makes contextual style decisions difficult to supervise, while the entanglement of style, timbre, and content often leads to weak instruction-following and severe timbre drift across turns. To overcome these limitations, we propose Interactive TTS, a dynamic, style-adaptive framework for contextually appropriate and speaker-consistent speech generation. Interactive TTS decouples the process by explicitly modeling contextual style decisions as executable instructions. To bridge the gap between style decisions and speech generation, we introduce Iterative Rejection Sampling Fine-Tuning (Iterative RSFT) and Context-Aware Direct Preference Optimization (CADPO), which significantly enhance instruction-following and align the generated speech with conversational contexts. Extensive experiments demonstrate that Interactive TTS outperforms state-of-the-art models on VStyle and SpeechParaling-Bench. Demo is available at https://wjtian-wonderful.github.io/InteractiveTTS/
arXiv ID: 2609.25707 / 要約の誤りについて