話しながら聞く対話AIのための合成会話データ
A Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations
この論文をやさしく読む
ひとことで言うと
話しながら聞く対話AIに、割り込みや相づちを区別して学ばせる合成音声データの作り方です。
何に役立つ?
全二重の音声対話で、いつ話し始め、聞き、発言権を譲るかを学習する際に役立ちます。
この研究の面白いところ
42種類の会話現象を作り、追加学習後のMoshiでは参照会話の発言を受け取る割合が0.44から0.85へ上がりました。
どこまで分かった?
結果は英語・中国語の合成会話と要旨にある評価条件によります。自然会話全般での効果は要旨からは判断できません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
発話中も相手の声を聞く全二重の対話システムは、発話の終了と発話途中の間を区別し、発言権を求める割り込みと短い相づちや第三者に向けた発話も区別する必要がある。しかし既存の会話コーパスでは、こうした出来事を十分に制御できず、意図のラベルも限られている。この研究は、出来事間の関係を記した一覧から、意図ラベル付きの二つの音声チャネルを持つ会話を合成するパイプラインを示す。大規模言語モデルが各出来事の話者、文章、会話行為、先行する出来事へのつながりを作成するが、絶対時刻は予測しない。各出来事を個別に音声合成し、元の文章に位置合わせして共通の時間軸に置く。そのため発話交代の時点は生成された音声から測定され、無音時間は指定するか発話交代の分布から抽出する。 パイプラインは英語と中国語で、八つの系列に属する42種類の現象を扱う。作成した意図からフレーム単位のシステム行動を導き、少数で多様な先行例と、自己申告の確率を伴う代案を求める一括プロンプトにより多様性を高める。要素を取り除く比較では、狙った多様性の各側面で改善が見られた。発言権を取る、保つ、譲る、持たないという四行動のラベル空間では、現在と過去の音声だけを使う意味的音声活動検出器が、話し始めと聞き始めのF1スコア0.819と0.802を達成した。自分で応答を生成する場合、全二重音声モデルMoshiが参照会話の発言を受け取る割合は、合成コーパスで追加学習する前の0.44から後には0.85になった。システムが発言権を持つかをフレーム単位で予測する適合率は0.46から0.88に上がった。各段階で参照文脈を与えた場合の発言権に関するF1は0.893から0.962に上がった。これらの結果は、制御された合成が全二重の発話交代管理に、学習可能で転移できる教師情報を提供しうることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited control over these events and limited labels for their intent. We present a pipeline for synthesizing intent-labeled, two-channel conversational speech from relational event lists. An LLM authors each event's speaker, text, conversational act, and attachment to an earlier event without predicting absolute timestamps. Events are synthesized independently, aligned with their source text, and placed on a shared clock, so turn-taking landmarks are measured from the rendered signal while silence durations are specified or sampled from turn-taking distributions. The pipeline covers 42 phenomena across eight families in English and Mandarin, derives frame-level system actions from authored intent, and promotes diversity using small, diverse sets of prior examples and batch prompts that request alternatives with self-reported probabilities. Ablations show gains in each targeted diversity dimension. On a four-action label space for taking, holding, releasing, and not holding the conversational floor, a semantic voice-activity detector using only current and past audio reaches start-speaking and start-listening F1 scores of 0.819 and 0.802. When generating its own responses, the full-duplex speech model Moshi takes 0.85 of the reference turns after fine-tuning on the generated corpus, compared with 0.44 before fine-tuning. Its frame-level precision for predicting system-floor occupancy rises from 0.46 to 0.88. With reference context at each step, its frame-level floor F1 rises from 0.893 to 0.962. These results show that controlled synthesis can provide learnable and transferable supervision for full-duplex turn management.
著者のコメント
Short version submitted to ICASSP 2027
arXiv ID: 2609.28806 / 要約の誤りについて