音声で数学問題を解くモデルを強化学習で改善
Voice of Reason: Reinforcement Learning for Spoken Math
この論文をやさしく読む
ひとことで言うと
音声で質問を受けて数学問題を解くモデルに、答えの検証に基づく強化学習を加え、正解率を高める研究です。
何に役立つ?
音声対話で計算や文章題に答えるシステムの学習方法として参考になります。音声モデルでも検証可能な報酬を使う効果を調べています。
この研究の面白いところ
追加の推論トークンなしでも改善し、さらにストリーミング推論と組み合わせると74.8%に達するという、二段階の効果を示しています。
どこまで分かった?
74.8%はGSM8Kで既存のストリーミング推論を組み合わせた条件の自由形式回答精度です。一般的な会話やあらゆる数学課題で同じ精度が得られるという結果ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声言語モデルは、カスケード型システムより豊かな人と機械の音声対話を可能にし、言葉以外の音声情報を利用でき、遅延も小さい。しかし、数学的推論のベンチマークでの精度は、テキストモデルに後れを取ってきた。検証可能な報酬を用いた強化学習(RL)は、複雑な問題を解くテキストモデルの能力の拡張と、幻覚の抑制に重要な役割を果たしてきた。 本研究では、テキストと音声による数学問題解決の隔たりを埋めるため、音声モデルGLM-4-Voice(Zengら、2024)へのRLの適用を検討する。まず、合成した音声の質問応答データによる教師あり微調整で、モデルを対象領域に適応させる。その後、追加の推論トークンがなくても、RLによってGSM8Kの精度が、従来の音声モデルでは補助的な推論過程を伴う場合にのみ達成されていた水準を上回ることを示す。 既存のストリーミング推論技術と組み合わせると、自由形式回答の正解率はさらに74.8%まで向上する。これは、音声を直接扱うモデルの、音声による数学能力において新たな最高性能となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Speech language models enable richer spoken interactions between humans and machines than cascaded systems, allowing access to paralinguistic information and lower latency. However, their accuracy on mathematical reasoning benchmarks has lagged behind those of text models. Reinforcement learning (RL) with verifiable rewards has been instrumental in extending text models' capabilities for solving complex problems and limiting hallucinations. In this work, we explore applying RL to the GLM-4-Voice speech model (Zeng et al., 2024) to bridge the gap between textual and spoken mathematical problem solving. We first adapt the model to the domain using supervised fine-tuning on synthesized spoken question-answering data. We then show that, even without extra reasoning tokens, RL improves the accuracy on GSM8K beyond levels previously achieved for speech models only with supplementary reasoning traces. When combined with existing streaming reasoning techniques, we show further gains to 74.8% free-form accuracy. This establishes a new state-of-the-art for mathematical spoken abilities with speech-native models.
著者のコメント
Accepted at COLM 2026
arXiv ID: 2609.18677 / 要約の誤りについて