arXiv論文メモ
新着一覧
eess.AS · 査読状況未確認

会議全体の話者分離における音声言語モデルの能力と限界

Speech Language Models for Full-Meeting Speaker Diarization: Capabilities and Limitations

Jialu Li, Jinchuan Tian, Shinji Watanabe

この論文をやさしく読む

ひとことで言うと

会議全体の音声で誰がいつ話したかを、音声言語モデルで推定する研究。

何に役立つ?

話者分離を音声認識の精度と切り離して評価・改善する参考になる。

この研究の面白いところ

開始・終了時刻を直接出す方式が、フレーム単位で出す方式より安定した点。

どこまで分かった?

会議全体では話者追跡と重複発話の見落としが残り、後処理による結びつけが必要だった。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルと音声基盤モデルを統合した音声言語モデルの進歩により、音声処理課題を統一的な系列モデルで扱えるようになった。しかし音声言語モデルを使う話者分離の多くは自動音声認識と密接に結びつき、単語単位の指標で評価されるため、認識精度と独立に話者分離の性能を測りにくい。本研究は、ESPnet-SpeechLMをトークンに基づいて話者分離の仮説を生成する基盤として調べ、音響入力を条件とする構造化トークンの自己回帰生成として話者分離を定式化する。話者交代の開始・終了時刻を明示する事象型表現と、フレーム単位の話者活動を予測するフレーム型表現を体系的に比較した。会話の構造化された手掛かりを与えるため、発話活動検出、発話の重なり検出、話者交代の計数も補助課題として出力系列に組み込んだ。複数の会議データセットでは、事象型表現がフレーム型より安定して一貫した話者分離出力を生んだ。音声言語モデルの出力は有用な時間構造を表していたが、会議全体での話者分離は録音を通じた話者追跡と重なった発話の見落としに制約されていた。明示的に話者を結びつける後処理は話者の取り違えを大きく減らし、頑健な話者分離には持続的な追跡と重なりを考慮した生成が必要だと示唆する。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recent advances in Speech Language Models (SpeechLMs), which integrate large language models with speech foundation models, have enabled unified sequence modeling of speech processing tasks. However, many SpeechLM-based approaches to speaker diarization (SD) are tightly coupled with automatic speech recognition (ASR) and evaluated using word-level metrics, making it difficult to assess SD performance independent of ASR accuracy. In this work, we investigate ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input. We systematically compare two output representations: an event-based representation that explicitly models speaker turn onset and offset timestamps, and a frame-based representation that predicts frame-level speaker activity. To provide structured conversational cues, we further incorporate auxiliary tasks including speech activity detection, overlapped speech detection, and speaker turn counting within the output sequence. Across multiple meeting datasets, we find that event-based representations produce more stable and consistent SD outputs than frame-based representations. Our analysis shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recording-level speaker tracking and overlap-related misses. Explicit speaker-linking post-processing substantially reduces speaker confusion, suggesting that robust SpeechLM-based SD requires persistent speaker tracking and overlap-aware generation.

著者のコメント

Accepted to IEEE SLT 2026

arXiv ID: 2609.23114 / 要約の誤りについて