合成会話で学ぶ音声アシスタントの文脈に応じた起動
Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations
この論文をやさしく読む
ひとことで言うと
起動語の後の発話がアシスタントへの命令かどうかを、会話の文脈から判定する研究。
何に役立つ?
考えられる用途は、音声アシスタントが周囲の無関係な会話へ反応するのを減らすこと。
この研究の面白いところ
直接の呼びかけに加え、続きの発話や宛先の違う発話を含む62.3時間の制御可能な合成会話を作った。
どこまで分かった?
効果は合成会話の複数条件で示されたもので、実際の家庭や職場での性能は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
起動語の検出は、利用者と音声アシスタントが円滑にやり取りするための入り口となる重要な機能である。本論文では、従来の起動語そのものの検出に、文脈に応じた起動判定を加える新しい起動システムを紹介する。最初に起動語が検出された後、システムは推論を使って利用者の命令と関係のない発話を区別し、効率的で文脈に合った応答を目指す。直接呼びかける発話、文脈上の続きの発話、アシスタント宛てではない発話を含む、複数話者の制御可能な会話を生成するデータ構築の仕組みを示し、62.3時間のコーパスを作成した。実験では、多様な合成会話の条件で提案方法の有効性を示した。再現性と音声アシスタント技術の研究を進めるため、コード、データセット、学習済みモデルを公開する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Wake word detection is a critical component of virtual assistants, serving as the gateway to seamless user interactions. This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection. After an initial wake word activation, the system uses reasoning to distinguish between user commands and unrelated speech, ensuring efficient and context-aware engagement. We present a data generation architecture that produces a 62.3-hour corpus of controllable multi-speaker conversations containing direct invocations, contextual follow-ups, and non-addressed speech. Experimental results demonstrate the effectiveness of the proposed approach across diverse synthetic conversational scenarios. We release the code, dataset and trained models to promote reproducibility and further advancements in intelligent assistant technologies.
arXiv ID: 2609.27037 / 要約の誤りについて