単一マイク録音から音源と距離の順序を推定
DAMSEP: Distance-Aware Monaural Source Separation using Multi-RIR Estimation
この論文をやさしく読む
ひとことで言うと
一つのマイクで録った混合音から、音を分けると同時に各音源の相対的な遠近を推定する方法。
何に役立つ?
単一マイク録音で音源の内容と空間的な手掛かりを一緒に扱う手法の評価に役立つ。
この研究の面白いところ
音源ごとの部屋の響きを推定して直接音と残響音の比から遠近を求め、専用データセットで複数の性能を検証した。
どこまで分かった?
距離は相対的な近い・遠いの順序として扱う。主な評価はシミュレーション条件で、実測RIRによる追加評価の範囲は要旨に詳述されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
部屋のインパルス応答(RIR)には音源までの距離の手掛かりが含まれるが、従来の単一マイクによる音源分離は音声内容の復元を中心とし、音源ごとのRIRを推定しないため、関連する空間情報を失う。この制約に対し、著者らは単一マイクの混合音から音源分離と複数音源のRIR推定を同時に学習する初のエンドツーエンド枠組み、DAMSEPを提案する。DAMSEPは、分離の基幹モデルに共通の残響除去モジュールとRIR推定モジュールを組み合わせ、音源推定と残響を含む音の再構成を目的として、残響のない各音源と音源ごとの複素畳み込み伝達関数を同時に復元する。対応するRIRの直接音と残響音の比から、音源の相対的な近い・遠いの順序も求められる。包括的な評価のため、異種の音源内容、多様なシミュレーション室内条件、音源ごとのRIR、幾何学的な距離の注釈を含むHETMIXRを導入する。HETMIXRでの実験では、音源分離、RIR推定、距離順序のいずれでも優れた性能を示した。構成要素を除く実験は音源の教師信号と残響音の再構成が相補的に役立つことを示し、追加評価では単一話者入力と、未見の部屋で測定したRIRから生成した混合音への一般化も確認した。コードとデータセットは公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Although room impulse responses (RIRs) encode source-distance cues, conventional monaural source separation focuses on recovering audio content without estimating source-specific RIRs, losing the associated spatial information. To address this limitation, we propose Distance-Aware Monaural Source Separation using Multi-RIR Estimation (DAMSEP), the first end-to-end framework that is jointly trained for source separation and multi-source RIR estimation from a single-microphone mixture. DAMSEP integrates a separation backbone with shared dereverberation and RIR estimation modules to jointly recover clean sources and source-specific complex convolutive transfer functions under source estimation and reverberant reconstruction objectives, enabling relative near/far ordering through the direct-to-reverberant ratios of the corresponding RIRs. For comprehensive evaluation, we introduce HETMIXR, which spans heterogeneous source content and diverse simulated room conditions with source-specific RIRs and geometric distance annotations. Experiments on HETMIXR demonstrate superior performance in source separation, RIR estimation, and distance ordering. Ablation studies reveal the complementary benefits of source supervision and reverberant reconstruction, while additional evaluations show generalization to single-speaker inputs and mixtures generated using measured RIRs from an unseen room. Our code and dataset are available at https://github.com/Wenanzhi/DAMSEP.
著者のコメント
5 pages, 1 figure, 5 tables
arXiv ID: 2609.29749 / 要約の誤りについて