arXiv論文メモ
新着一覧
eess.AS · 査読状況未確認

音声と言語以外の音を共通に学ぶJASPER

JASPER: Joint Audio and Speech Pre-trained Encoder Representations

Geeth George, Ameenudeen P E, Hrishikesh H Pillai, Sriram Ganapathy

この論文をやさしく読む

ひとことで言うと

発話、一般の音、音楽を共通に扱うため、時間と周波数の両方を学ぶ音響エンコーダー。

何に役立つ?

複数種類の音を扱う下流課題で共通の表現を使う研究に役立つ。

この研究の面白いところ

発話向け事前学習モデルに長い音区間の時間・周波数予測を加える。

どこまで分かった?

要旨は比較基準を上回ったと述べるが、個別の課題名や性能差の数値は示していない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

音声と一般の音に対する自己教師あり学習は、主に別々の方向で発展してきた。発話モデルは時間領域の予測を重視し、一般の音の表現学習は時間と周波数のパターンに注目してきた。この分離は互換性の隔たりを生み、領域をまたぐ汎化を制限する。本研究は、発話で事前学習したモデルへ時間・周波数の目標を加える枠組み、Joint Audio and Speech Pre-trained Encoder Representations(JASPER)を導入する。具体的には、長い音の区間に対し、時間的・スペクトル的な目標のマスク付き予測を行い、発話と一般の音の時間・周波数表現を学ぶ。提案法は、発話、一般の音、音楽にわたる多様な課題で、複数の比較基準と既存の音声・音響エンコーダーを一貫して上回り、統一的な時間・周波数モデル化の有効性を示した。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Self-supervised learning (SSL) for speech and audio has largely progressed along separate tracks: speech models emphasise time-domain prediction, whereas audio representation learning has focused on time-frequency patterns. This separation creates a compatibility gap, limiting cross-domain generalization. In this work, we introduce JASPER, Joint Audio and Speech Pre-trained Encoder Representations, a framework that augments speech-pretrained models with time-frequency objectives. Specifically, JASPER performs masked prediction of temporal and spectral targets over long audio segments, enabling spectro-temporal representation learning of speech and audio signals. The proposed method consistently outperforms multiple baselines and existing speech/audio encoders on diverse speech, audio, and music tasks, demonstrating the effectiveness of unified spectro-temporal modeling.

arXiv ID: 2609.27260 / 要約の誤りについて