arXiv論文メモ
新着一覧
eess.AS / eess.SP · 査読状況未確認

残響のある立体音響を生成モデルで高解像度化する

Generative Learning for Ambisonic Upscaling

Amit Milstein, Nir Shlezinger, Boaz Rafaely

この論文をやさしく読む

ひとことで言うと

残響のある録音から、音がどの方向から来るかをより細かく表現する立体音響成分を生成します。

何に役立つ?

考えられる用途は、低次アンビソニックス録音の空間表現の改善です。数値評価に加え、人が聴いた品質と方向の正確さも調べています。

この研究の面白いところ

残響で単純な方向推定が難しくなることに対し、1つの決定的な復元ではなく生成モデルを使って対応しています。

どこまで分かった?

優位性は評価した残響条件内の結果です。要旨には聴取者数、評価尺度の値、計算費用などの具体値はありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

アンビソニックスの高解像度化(AU)は、低次の観測から高次アンビソニックス(HOA)成分を推定し、音場の空間解像度を高めることを目指す。深層学習とモデルベースの手法がAUに検討されてきたが、現実的な状況ではどちらも性能が大きく低下する。残響のある音場が、識別的な対応付けに内在する方向の疎性の前提を破るためである。 本研究ではAUを決定論的な再構成ではなく生成課題として扱い、残響を含む音声の空間情報の復元を特に対象として生成モデルを拡張する。代表的な2つの連続時間生成手法を調べ、スコアベース生成モデルとフローマッチングの両方を、こうした複雑な音響設定に適応させる。さまざまな音響条件で最先端の基準手法と比較する広範な数値研究を行う。さらに、複数の残響条件で、提案する生成枠組みの知覚品質と空間的な正確さを評価する主観的な聴取試験を実施する。これらの研究は、フローマッチングが、すべての評価した残響条件で識別的な手法と拡散ベースの手法の両方を一貫して上回ることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Ambisonics Upscaling (AU) aims to enhance the spatial resolution of sound fields by estimating high-order Ambisonics (HOA) components from low-order observations. While deep learning and model-based strategies have been considered for AU, both approaches exhibit significant performance degradation in realistic scenarios, where reverberant sound fields violate the directional sparsity inherent to discriminative mappings. In this work, we address AU as a generative task rather than a deterministic reconstruction, expanding generative modeling to specifically target the recovery of spatial information in reverberant speech. We investigate two dominant continuous-time generative paradigms, adapting both Score-based Generative Model and Flow Matching to these complex acoustic settings. We provide an extensive numerical study comparing our methods against state-of-the-art baselines in various acoustic scenarios. Additionally, we conduct subjective listening tests to evaluate the perceived quality and spatial accuracy of the proposed generative framework across various reverberant scenarios. The studies reveal that Flow Matching consistently outperforms both its discriminative counterparts and Diffusion-based paradigms in all reverberant settings.

arXiv ID: 2609.23479 / 要約の誤りについて