arXiv論文メモ
新着一覧
cs.CL / cs.SD / eess.AS · 査読状況未確認

音声モデルが推論中の無音を減らす途中発話

Spoken Language Models that Think Aloud

Junyi Ao, Kainan Peng, Mingbo Ma, Shun Zhang, Zhenyu Tang, Xutai Ma, Xiang Li, Yinghao Li, Yuancheng Wang, Zhizheng Wu, Haizhou Li, Qing He, Xubo Liu

この論文をやさしく読む

ひとことで言うと

音声モデルが回答を考えている間、短い進捗発話を挟んで無音を減らす方法。

何に役立つ?

音声対話で長い待ち時間を感じさせにくいシステム設計に役立つ可能性がある。

この研究の面白いところ

推論と途中発話を別の流れで進め、最終回答ができたら不要な途中発話を取り消す。

どこまで分かった?

要旨が示すのはベンチマーク上の無音時間と回答精度であり、実際の利用者体験全般の評価は示されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

思考の連鎖(CoT)による推論は言語モデルの能力を高めたが、音声言語モデルにそのまま適用すると、先に考えてから話す逐次方式では長い無音時間が生じ、リアルタイムの会話を妨げる可能性がある。この問題に対し、Thinker–Talker構成の推論型音声言語モデル向けに、非同期の途中発話の枠組みを提案する。論理的な推論を行う主たる流れと、ユーザー入力および変化する推論状態に応じて、課題に即した短い進捗発話を生成する軽量な流れを維持する。動的な調整方式が実行時に両者を連携させ、無音を避けるために追加の途中発話を起動し、最終回答が準備できたときには待機中の発話を取り消す。音声推論と質疑応答のベンチマーク実験では、逐次的な先に考えてから話す基準方式と同程度の回答精度を保ちつつ、推論中にユーザーが聞く無音時間を大幅に減らした。非同期の途中発話が音声言語モデルの応答性を高め得ることを示している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

While Chain-of-Thought (CoT) reasoning has improved the capability of language models, directly applying it to Spoken Language Models (SLMs) may introduce long silent intervals under the serial "think-then-speak" paradigm, disrupting real-time spoken interaction. To address this issue, we propose an asynchronous think-aloud framework for reasoning-based SLMs within the Thinker-Talker architecture. The framework maintains a primary reasoning stream for logical deduction and a lightweight think-aloud stream that generates short, task-grounded progress utterances conditioned on the user input and the evolving reasoning state. A dynamic balance strategy coordinates the two streams at runtime, triggering additional think-aloud speech to avoid silent gaps and canceling pending utterances when the final response becomes ready. Experiments on spoken reasoning and question-answering benchmarks show that our approach substantially reduces user-audible silence during reasoning while maintaining answer accuracy comparable to that of a serial "think-then-speak" baseline, demonstrating the potential of asynchronous think-aloud for responsive interaction in SLMs.

著者のコメント

Accepted at SLT 2026

arXiv ID: 2609.26488 / 要約の誤りについて