音声情報を根拠にした推論をテスト時学習で促す
Listen Then Reason: Perception-Grounded Test-Time Reinforcement Learning for Large Audio-Language Models
この論文をやさしく読む
ひとことで言うと
音声を聞くモデルが、言葉だけの推測に寄り過ぎず、入力された音を推論の根拠に使うようテスト時に学習する方法です。
何に役立つ?
正解ラベルのないテスト音声を使って、音声推論モデルの適応を行うための手法です。音声情報をどの層でどれだけ使っているかを診断する見方も提供します。
この研究の面白いところ
まず音響情報への依存と性能の関連を分析し、その観点をテスト時強化学習の設計につなげます。単に推論を長くするのではなく、音声への根拠付けを重視します。
どこまで分かった?
音響依存度と精度の分析は関連を示したもので、そこだけで因果関係が確立するわけではありません。複数モデルで改善を報告していますが、改善幅やテスト時の計算負荷は要旨にありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模音声言語モデル(LALM)は、より幅広い音声推論タスクに利用されつつある。これらのモデルは通常、音声表現を大規模言語モデル(LLM)の基盤に組み込み、マルチモーダル推論を可能にする。最近のテスト時強化学習(TTRL)は、事前学習後にラベルなしのテストデータを活用し、LLMの推論能力をさらに改善する。しかし、LALMの知覚能力の重要性、とりわけ推論中に音響的な証拠をどれほど取り込み、依拠し、それが最終的なタスク性能にどう寄与するかは、十分に調べられていない。この不足は、音声推論向けのTTRLのような効果的な事後学習法の開発を制限する。本研究ではまず、推論過程で音声情報がどのように統合・利用されるかを解析する。層ごとの知覚への依存度を定量化し、音響情報への依存が強いほど、正解率が高く、音声入力に帰せられる性能向上も大きいという関連を示す。これを基に、ラベルを使わないテスト時の最適化を、知覚情報に根拠を置く推論に沿わせるPerception-Grounded TTRL(PG-TTRL)を提案し、モデルが音声入力をより強く基礎とした推論を組み立てるよう促す。複数のLALMとベンチマークでの実験により、PG-TTRLは基礎モデルと標準的なTTRLの双方より推論性能を一貫して改善し、テスト時の音声推論で知覚的根拠を重視して最適化する価値を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-20(UTC)
- 最新改訂
- 2026-09-20 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-20 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio representations into a large language model (LLM) backbone to enable multimodal reasoning. Recent test-time reinforcement learning (TTRL) methods further improve LLM reasoning capability by leveraging unlabelled test data after pre-training. However, the importance of the perceptual capability of LALMs remains underexplored, particularly how much acoustic evidence is integrated and relied upon during reasoning, and how this contributes to final task performance. This gap limits the development of effective post-training methods like TTRL for audio reasoning. In this work, we first analyse how audio information is integrated and utilised during reasoning process. We quantify layer-wise perceptual reliance and show that stronger acoustic reliance is associated with higher accuracy and a larger performance gain attributable to the audio input. Building on this, we propose Perception-Grounded TTRL (PG-TTRL), which aligns label-free test-time optimisation with perceptually grounded reasoning, encouraging the model to structure its reasoning more strongly on the audio input. Experiments across LALMs and benchmarks show that PG-TTRL consistently improves reasoning performance over both the base models and standard TTRL, showing the value of perceptual-grounding optimisation for test-time audio reasoning.
arXiv ID: 2609.23589 / 要約の誤りについて