話者や感情の情報で音声言語モデルを事前学習する
Teaching LLMs to Hear Who Spoke What: Metadata-Supervised Pretraining for Encoder-Free Speech-LLMs
この論文をやさしく読む
ひとことで言うと
音声の内容だけでなく、誰が話したかや感情の情報を教師信号として使い、専用音声エンコーダーなしで言語モデルに音声能力を学ばせる研究です。
何に役立つ?
会話の書き起こしと話者の区別、言葉以外の音声情報の理解を一緒に扱うモデルを作るための学習方法になります。
この研究の面白いところ
事前学習済み音声エンコーダーが抽出した特徴ではなく、低水準の音響情報からLLMが学ぶ構成です。話者を意識した発話の組合せも学習に取り入れています。
どこまで分かった?
同量データでの比較相手はランダム初期化したエンコーダーありモデルです。大規模事前学習済みモデルとの比較では複数条件で上回るという結果で、すべての音声課題での優位を示してはいません。具体的な誤り率は要旨にありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
エンコーダーを使う音声大規模言語モデル(Speech-LLM)は一般に、言語内容を優先する事前学習済み音声エンコーダーを用いる。しかしそれによって、話者の識別やパラ言語的な理解に不可欠な細かな音響的手掛かりが捨てられる場合がある。対して、エンコーダーを使わないSpeech-LLMは、軽量な埋め込み層を通じてメルスペクトログラム特徴を直接LLMの入力空間へ写し、LLMが低水準の音響特徴から学べるようにする。しかし、このようなモデルの体系的な事前学習方法は十分に研究されておらず、大規模に事前学習した音声エンコーダーの不在を補う能力が制限されている。 本研究では、話者の識別情報や感情などの音声属性を利用し、話者を区別する能力とパラ言語的能力を育てる、メタデータを教師信号とする事前学習(MSP)を提案する。さらに、話者識別を強める話者を考慮した発話構成(SAUC)を導入し、ランダムな区間マスキングで事前学習を正則化する。主に複数話者の会話で、音声認識(ASR)と話者ダイアライゼーションを同時に行う課題を評価し、パラ言語的な音声理解課題の実験も補足する。学習データの条件をそろえた比較では、本研究のエンコーダーなしモデルは、ランダム初期化したエンコーダーありの対応モデルを上回る。メタデータ付きデータが限られていても、はるかに大規模なコーパスで事前学習した音声エンコーダーを使うモデルに対して競争力があり、複数の設定でそれらを上回る。これらの結果は、多様な音声能力を獲得するネイティブなマルチモーダルLLMを構築するうえで、エンコーダーなしの構造が持つ可能性を示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Encoder-based speech large language models (Speech-LLMs) commonly employ pretrained speech encoders that prioritize linguistic content but may discard fine-grained acoustic cues essential for speaker discrimination and paralinguistic understanding. Encoder-free Speech-LLMs instead map Mel-spectrogram features directly into the LLM input space through lightweight embedding layers, enabling the LLM to learn from low-level acoustic features. However, systematic pretraining strategies for encoder-free Speech-LLMs remain underexplored, limiting their ability to compensate for the absence of large-scale pretrained speech encoders. We propose metadata-supervised pretraining (MSP), which leverages speech attributes such as speaker identity and emotion to develop speaker-discriminative and paralinguistic capabilities. We further introduce speaker-aware utterance composition (SAUC) to strengthen speaker discrimination and apply random span masking to regularize pretraining. We primarily evaluate our approach on joint ASR and speaker diarization in multi-speaker conversations, complemented by experiments on paralinguistic speech-understanding tasks. Under matched training-data conditions, our encoder-free model outperforms its randomly initialized encoder-based counterpart. With limited metadata-annotated data, it is competitive with models using speech encoders pretrained on substantially larger corpora, outperforming them in several settings. These results demonstrate the potential of encoder-free architectures for building native multimodal LLMs that acquire diverse speech capabilities.
arXiv ID: 2610.01695 / 要約の誤りについて