arXiv論文メモ
新着一覧
cs.SD / cs.MM · 査読状況未確認

音声モデルの表現は音の足し合わせを捉えるか

Do Audio Representations Compose Additively?

Chenhao Xue, Zhijin Guo, Joyraj Chakraborty, Martin Reed, Nikolaos Thomos

この論文をやさしく読む

ひとことで言うと

音声モデルの内部表現で、複数の音源を足し合わせた場面を表せるか調べる研究。

何に役立つ?

考えられる用途は、音声表現が複雑な音の組み合わせを捉えるかの診断。

この研究の面白いところ

CLAPは大きい未知の組み合わせでも基準より良かったが、音声向けの二モデルは一部の基準を上回らなかった。

どこまで分かった?

三モデルとも再構築に残差があり、完全に加算的な表現ではない。文章との学習がCLAPの差の原因かは推測にとどまる。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

複雑な音の場面を単純な音源の組み合わせとして表す能力は、聴覚の知覚や古典的な加算型の信号モデルで中心的である。しかし現代の事前学習済み音声表現が、組み合わせ方を教わらなくても加算的な構造を内包するかは不明である。従来の評価は音と文章の対応に重点を置き、文章との結び付けに依存しない音声表現そのものの加算性は十分に調べられていない。 本研究は、固定した音声表現に二段階の診断を行う。まず正準相関分析で、表現と音源ラベルの線形な対応を測る。次に、各音源ラベルの正確な組み合わせで音声をまとめて表現を平均し、学習時に使った組み合わせだけから推定した音源ごとの寄与で、保留した組み合わせの平均を再構築できるか試す。FSD50KとCHiME-Homeで、Wav2Vec2、HuBERT、CLAPの表現を調べた。 三つのモデルは、ラベルを入れ替えた基準より線形相関が高く、未知の組み合わせの再構築も正確だった。より大きい組み合わせを保留すると、FSD50KではCLAPがラベル入れ替えやラベルの重なりだけによる基準を上回ったが、音声向けの二モデルはラベルの重なりによる基準を上回らなかった。コサイン類似度の改善が大きかったのもCLAPだけで、多様な音と文章で学習したことが関係する可能性がある。すべてのモデルに再構築の残差があり、非線形または組み合わせに分解できない音の構造など、加算的な表現の限界も示された。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Compositionality, the ability to represent complex acoustic scenes as combinations of simpler sound sources, is central to auditory perception and classical additive signal models. Still, it remains unclear whether modern pre-trained audio representations internalize additive structure without compositional supervision. Existing evaluation frameworks of audio compositional reasoning largely focus on cross-modal audio-text alignment, leaving open whether audio representations themselves exhibit additive compositional structure independent of text grounding, analogous to vector arithmetic in word representations. To investigate this, we adopt a two-step diagnostic for frozen audio representations. First, we quantify linear alignment between representations and sound source labels using canonical correlation analysis. Second, we test additive compositional generalization via leave-one-combination-out reconstruction, grouping clips by exact source-label set, averaging their representations, and predicting held-out means from per-source contributions fitted only on training combinations. With larger combination holdouts, CLAP outperforms the permuted and label-overlap baselines on FSD50K, while the speech models do not outperform the label-overlap baseline. We examine representations generated by Wav2Vec2, HuBERT, and CLAP on FSD50K and CHiME-Home datasets. All three models show consistently higher linear correlation and more accurate leave-one-combination-out reconstructions than the permuted baselines. However, only CLAP shows large cosine similarity gains, which could be associated with its training on many kinds of audio and text. Finally, we note that all three models exhibit reconstruction residuals, revealing limits of additive compositionality such as nonlinear or non-compositional audio structure.

著者のコメント

Submitted to IEEE Signal Processing Letter

arXiv ID: 2609.27187 / 要約の誤りについて