声の生成と話し方の理解を相互に学習するCycleSpeech
CycleSpeech: Reciprocal Alignment for Instruction-Controlled Speech Synthesis and Paralinguistic Understanding
この論文をやさしく読む
ひとことで言うと
指示どおりの話し方で音声を作るモデルと、その話し方を読み取るモデルを、共通の声の属性表を使って互いに学習させる方法です。
何に役立つ?
考えられる用途は、口調や話し方を指示できる音声合成と、音声の属性を取り出す解析です。人の好みの注釈を別途集めずに、両課題を結び付ける学習法を示しています。
この研究の面白いところ
生成して理解し直す方向と、理解して生成し直す方向を両方使います。学習中に相手モデルが変わっても、固定した目標プロファイルを基準にする設計です。
どこまで分かった?
評価言語は中国語と英語です。4.50と10.06は正解率のパーセントポイント差で、相対的な改善率ではありません。注釈不要というのは人の好みの注釈についてで、教師ありデータセットは使っています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
指示に従う音声合成と、話し方などのパラ言語情報の理解は、しばしば独立に学習され、両課題の相互フィードバックは十分に探究されていない。本研究では、教師信号と相互フィードバックの共通目標となる、共有の構造化された声のプロファイルを通じて、生成と理解を結ぶ枠組みCycleSpeechを導入する。 順方向のサイクルでは、生成音声から復元したプロファイルと目標プロファイルを比較し、意図した属性を合成音声が表しているかを評価する。逆方向のサイクルでは、実音声から推定したプロファイルが、元の話し方の再構成を導けるかを評価する。両方向を支えるため、指示、目標音声、話者参照、構造化プロファイルを対応付けた20,046例の2言語データセットを構築する。 共同の教師あり微調整を基に、CycleGRPOはプロファイルの整合性と話し方の再構成に根拠を置く相互報酬を用い、方策更新を交互に行う。固定された目標プロファイルが、変化し続ける相手モデルからのフィードバックの基準となる。この手順は、人の好みに関する注釈も、好みのデータで学習した追加の報酬モデルも必要としない。 中国語と英語のベンチマークでは、競争力のある合成品質を保ちながら、指示への適合とプロファイルの復元が改善した。Step-Audio-2-miniと比べ、CycleSpeechは指示一致正解率を中国語で4.50ポイント、英語で10.06ポイント改善する。条件を統制したアブレーション実験も、生成制御へのサイクル型フィードバックの寄与を支持する。これらの結果は、構造化された声のプロファイルが、音声生成とパラ言語理解の相互学習の接点になることを支持する。オンラインデモは https://cyclespeech.github.io で利用できる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Instruction-controlled speech synthesis and paralinguistic understanding are often trained independently, leaving reciprocal feedback between the two tasks underexplored. We introduce CycleSpeech, a framework that connects generation and understanding through a shared, structured voice profile that serves as a common target for supervision and reciprocal feedback. The forward cycle assesses whether synthesized speech expresses the intended attributes by comparing recovered and target profiles. The backward cycle evaluates whether profiles inferred from real speech can guide reconstruction of the source speaking style. To support both directions, we construct a bilingual dataset of 20,046 examples pairing instructions, target speech, speaker references, and structured profiles. Building on joint supervised fine-tuning, CycleGRPO alternates policy updates using reciprocal rewards grounded in profile consistency and speaking-style reconstruction. Fixed target profiles anchor feedback from the evolving counterpart. This procedure requires neither human preference annotations nor an additional preference-trained reward model. Evaluations on Chinese and English benchmarks show improved instruction adherence and profile recovery while maintaining competitive synthesis quality. Compared with Step-Audio-2-mini, CycleSpeech improves instruction-match accuracy by 4.50 and 10.06 percentage points in Chinese and English, respectively. Controlled ablations further support the contribution of cycle feedback to generation control. These results support structured voice profiles as an interface for reciprocal training between speech generation and paralinguistic understanding. An online demo is available at https://cyclespeech.github.io.
arXiv ID: 2609.24771 / 要約の誤りについて