音の特徴を学ぶデータとモデルを繰り返し改善するEvoAudio
EvoAudio: Recursive Self-Improvement for Audio Understanding
この論文をやさしく読む
ひとことで言うと
音声モデルの苦手な音の特徴に合わせ、音声や質問を作り直しながら学習を繰り返す手法です。
何に役立つ?
細かな音響注釈を人手で大量に作れない場合に、検証可能な学習問題を作りながら音声理解を改善する研究に役立つ。
この研究の面白いところ
現在のモデルの成績が次のデータの焦点と難易度を決め、音声の生成過程から正解を決められる質問を作る閉ループになっている。
どこまで分かった?
改善は五つのモデルと三つの評価課題、13回の反復で報告された。要旨には新たな人手注釈なしとあるが、すべての音声課題への一般化は示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声言語モデルは、音がどのように聞こえるかよりも、何が話されているかをよく理解する。この差を埋めるには、単にデータを増やすだけでは足りない。細かな音響注釈には費用がかかり、より強いモデルが付けたラベルにはそのモデルの誤りや限界が引き継がれ、固定されたデータは学習者が上達しても適応できない。そこで、音声理解のための再帰的な自己改善システムEvoAudioを提案する。著者らの知る限り、モデル、音声波形、質問、難易度を一つの閉じた循環で進化させる初めての方法である。EvoAudioは現在のモデルの成績を用いて、次の学習データで重点を置く内容と難易度を決める。音声ツール群が、音声の作り方から答えが導ける質問を構成し、新たな人手の注釈なしに検証可能な教師信号を与える。強化学習でモデルを更新し、検証によって次の進化段階に採用するかを決める。13回の反復を通じ、異なる音声エンコーダーと言語モデルを持つ五つのモデルが、MMSU、MMAU-Pro、MMARで改善した。すべての基本モデルで最高の平均成績を達成し、総合性能は最大6.3ポイント向上した。改善は各反復を通じて進み、強くなったモデルが次の反復の起点となる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-improvement system for audio understanding. To our knowledge, it is the first to evolve the model, waveforms, questions, and difficulty in one closed loop. EvoAudio uses the current model's performance to set the focus and difficulty of the next training data. A library of audio tools then constructs questions whose answers follow from how the audio was made, providing verifiable supervision without new human annotation. Reinforcement learning updates the model, and validation decides whether it enters the next evolution round. Across 13 rounds, EvoAudio improves five models with different audio encoders and language backbones on MMSU, MMAU-Pro, and MMAR. It achieves the highest average for every backbone, raising overall performance by up to 6.3 points. The improvement unfolds over successive rounds, with each stronger model starting the next round.
arXiv ID: 2609.27389 / 要約の誤りについて