arXiv論文メモ
新着一覧
cs.SD / cs.AI · 査読状況未確認

アンビソニックスから音場の位置情報を自己教師学習

Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics

Goksenin Yuksel, Marcel van Gerven, Kiki van der Heijden

この論文をやさしく読む

ひとことで言うと

音を聞き分ける既存モデルに、立体音響から学んだ音の方向情報を後付けする研究。

何に役立つ?

考えられる用途は、既存の音響エンコーダーを再学習せずに音イベントの位置推定機能を加えること。

この研究の面白いところ

ラベルのない一次アンビソニックスで空間表現を学び、固定した音響モデルには軽量な融合部だけを接続する。

どこまで分かった?

報告されたFスコアは指定データセットでの結果であり、他の録音環境での性能は要旨からは分からない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

近年の自己教師学習による音声エンコーダーは、音環境の汎用的な表現を学べるが、空間情報を捉えない。この不足を補うため、ラベルのない一次アンビソニックスから音場の埋め込み表現を学ぶ自己教師学習の枠組みBearingsを導入する。マスク付きオートエンコーダーを事前学習し、市販の単一チャンネル音声エンコーダーから得られる固定された音響埋め込みを条件とするデコーダーと組み合わせる。 得られた音場埋め込みは再利用できる情報の流れとなり、軽量で学習可能な融合部を使って、重みを固定した音響エンコーダーへ接続できる。音イベントの定位・検出では、この音場埋め込みを音響表現と連結すると不足していた空間情報が加わり、検出と定位を同時に行える。位置に依存するFスコアは、TAU-NIGENS 2021で4未満から50へ、STARSS23では39へ上がった。著者らの知る限り、どちらのモデルも再学習せずに、固定された音響エンコーダーへ埋め込みを差し込める自己教師学習の音場エンコーダーはBearingsが初めてである。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Recently proposed self-supervised audio encoders learn powerful general-purpose representations of sound scenes, yet they are spatially blind. To supply the missing spatial representation of sound scenes, we introduce Bearings. Bearings is a self-supervised framework that learns soundfield embeddings from unlabeled first-order Ambisonics. We pre-train a masked auto-encoder paired with a decoder conditioned on frozen acoustic embeddings from an off-the-shelf single-channel audio encoder. Our results show that the resulting soundfield embeddings form a reusable stream that can be attached to frozen acoustic encoders with a lightweight trainable fusion head. On sound event localization and detection, concatenating our soundfield embeddings with acoustic representations provides the missing spatial information and enables joint detection and localization, raising the location-dependent F-score from below 4 to 50 on TAU-NIGENS 2021 and 39 on STARSS23. To our knowledge, Bearings is the first self-supervised soundfield encoder whose embeddings plug into frozen acoustic encoders without retraining either model.

arXiv ID: 2609.23152 / 要約の誤りについて