arXiv論文メモ
新着一覧
cs.CL / cs.LG / cs.SD · 査読状況未確認

音声AIの内部推論を調整して応答の待ち時間を減らす

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu, Shengbo Cai, Qinke Ni, Wan Lin, Tao Feng, Yingda shen, Ming-Hao Hsu, Zhixian Zhao, Liqiang Zhang, Teddy Sun, Steve Yves, Zhizheng Wu

この論文をやさしく読む

ひとことで言うと

音声AIが文章で思考過程を出力する代わりに内部表現で推論し、質問の難しさに応じて推論量を変える方式です。

何に役立つ?

音声を理解して考える性能を保ちつつ、回答を待つ時間を減らす用途に役立ちます。最初の回答トークンまでの時間を測っており、音声出力が完了するまでの総時間とは区別が必要です。

この研究の面白いところ

複数の推論候補を内部の分布で扱い、先の状態をまとめて予測します。難問では推論を増やす一方、簡単な問題では短く済ませる仕組みを学習しています。

どこまで分かった?

2種類の基盤モデルでの評価です。Qwen2.5-Omniで0.10秒を報告していますが、直接回答の0.05秒よりは長くなっています。11.8倍という値と表示時間は原要旨の表記を保持しています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

音声言語モデルとの対話の質は、モデルの知的能力と応答速度の両方に左右されるが、これらの両立は難しい。明示的な思考の連鎖(CoT)は推論と音声理解を改善するものの、中間の推論トークンの生成が応答を遅らせる。細かな音響的手掛かりを記述すると、CoTがさらに長くなり遅延も増える。潜在空間での推論はこの負担を減らせるが、既存手法はCoTに劣ることが多く、単一経路の教師信号や、問題の難しさに適応しない推論予算に制約されている。 本研究はAURALを提案する。潜在空間で複数のありうる推論の続きを分布として表現し、将来の状態をまとまりとして共同予測することで、逐次的な順伝播の回数と推論遅延を減らす。潜在推論に初期の教師信号を与えるため、感情認識、共感的な対話、一般的な推論に関する簡潔なCoTを備えた、2言語の音声発話68万3,000件、約1,000時間からなるAuralReason-683Kを構築した。AURAL-RLはこれらの推論記録を超えて探索し、質の高い回答につながる簡潔な推論に報酬を与え、問題ごとに推論量を調整する。 2種類の基盤モデルで、AURAL-RLはCoT-RLと同程度の性能を達成し、大半の指標で、それぞれの教師あり学習済み時点からの改善がより大きかった。さらに分析により、難しい質問ほど潜在推論のステップ数が増えることを示した。Qwen2.5-Omniでは、最初の回答トークンまでの時間を1.22秒から0.10秒へ、11.8倍高速化した。直接回答する場合は0.05秒である。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-10-01(UTC)
最新改訂
2026-10-01 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing fine-grained acoustic cues further lengthens CoT and increases latency. Latent reasoning can reduce this overhead, yet existing methods often trail CoT and remain limited by single-path supervision and reasoning budgets that do not adapt to problem difficulty. We introduce AURAL, which models a distribution over multiple plausible reasoning continuations in latent space and jointly predicts chunks of future states to reduce sequential forward passes and reasoning latency. To provide initial supervision for latent reasoning, we construct AuralReason-683K: 683K bilingual speech utterances (about 1,000 hours) with concise CoT for emotion recognition, empathetic dialogue, and general reasoning. AURAL-RL then explores beyond these traces, rewarding concise reasoning that yields high-quality answers and adapting reasoning effort to each problem. Across two backbones, AURAL-RL achieves performance comparable to CoT-RL, with larger gains over the respective supervised checkpoints on most metrics. Analysis further shows that harder questions elicit more latent reasoning steps. On Qwen2.5-Omni, it reduces time to the first answer token by 11.8x, from 1.22 to 0.10 s, versus 0.05 s for direct answering.

arXiv ID: 2610.01560 / 要約の誤りについて