音声と言語以外の音を共通に学ぶJASPER
JASPER: Joint Audio and Speech Pre-trained Encoder Representations
この論文をやさしく読む
ひとことで言うと
発話、一般の音、音楽を共通に扱うため、時間と周波数の両方を学ぶ音響エンコーダー。
何に役立つ?
複数種類の音を扱う下流課題で共通の表現を使う研究に役立つ。
この研究の面白いところ
発話向け事前学習モデルに長い音区間の時間・周波数予測を加える。
どこまで分かった?
要旨は比較基準を上回ったと述べるが、個別の課題名や性能差の数値は示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音声と一般の音に対する自己教師あり学習は、主に別々の方向で発展してきた。発話モデルは時間領域の予測を重視し、一般の音の表現学習は時間と周波数のパターンに注目してきた。この分離は互換性の隔たりを生み、領域をまたぐ汎化を制限する。本研究は、発話で事前学習したモデルへ時間・周波数の目標を加える枠組み、Joint Audio and Speech Pre-trained Encoder Representations(JASPER)を導入する。具体的には、長い音の区間に対し、時間的・スペクトル的な目標のマスク付き予測を行い、発話と一般の音の時間・周波数表現を学ぶ。提案法は、発話、一般の音、音楽にわたる多様な課題で、複数の比較基準と既存の音声・音響エンコーダーを一貫して上回り、統一的な時間・周波数モデル化の有効性を示した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Self-supervised learning (SSL) for speech and audio has largely progressed along separate tracks: speech models emphasise time-domain prediction, whereas audio representation learning has focused on time-frequency patterns. This separation creates a compatibility gap, limiting cross-domain generalization. In this work, we introduce JASPER, Joint Audio and Speech Pre-trained Encoder Representations, a framework that augments speech-pretrained models with time-frequency objectives. Specifically, JASPER performs masked prediction of temporal and spectral targets over long audio segments, enabling spectro-temporal representation learning of speech and audio signals. The proposed method consistently outperforms multiple baselines and existing speech/audio encoders on diverse speech, audio, and music tasks, demonstrating the effectiveness of unified spectro-temporal modeling.
arXiv ID: 2609.27260 / 要約の誤りについて