測定方向が異なっても使える頭部伝達関数の補間
HRTF Upsampling Across Varying Measurement Configurations with Geometry-Aware Query-Conditioned Aggregation
この論文をやさしく読む
ひとことで言うと
少数の方向で測った個人のHRTFから、別の方向の値を補い、測定方向の組合せが変わっても同じモデルを使う方法。
何に役立つ?
個人向け空間音響を作る際に、必要なHRTF測定点を減らす用途が考えられる。
この研究の面白いところ
目標と測定方向の相対的な幾何をクロスアテンションに入れ、一つのモデルで四つの標準配置と未学習の配置を扱った。
どこまで分かった?
最低の対数スペクトル歪みはSONICOMデータセットでの比較結果である。実際の聴感評価やすべての測定配置での性能は要旨に示されていない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
個人に合わせた頭部伝達関数(HRTF)は空間音響の再生に欠かせないが、一人ひとりについて多くの方向で測定するには時間と費用がかかる。HRTFの高密度化は、少数の測定から多数の方向のHRTFを推定して、この負担を減らす。近年の学習法は有望な性能を示しているものの、あらかじめ決めた測定方向の組合せに縛られる手法が多い。本研究は、測定配置が異なっても一つの学習済みモデルを使える、可変の入力測定に対応したHRTF高密度化手法GeoAttを提案する。GeoAttは、利用できる測定に対し、目標方向を条件とした幾何学的な空間情報の集約を周波数ビンごとに独立に行い、その後Conformerブロックで周波数領域をモデル化する。目標方向と測定方向の相対的な幾何情報は、クロスアテンションの加算バイアスとして組み込む。SONICOMデータセットでの実験では、一つの学習済みモデルが、Listener Acoustic Personalization(LAP)チャレンジの標準的な四つの測定配置すべてで最小の対数スペクトル歪みを達成し、学習時に明示的には含めなかった配置にも一般化した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Personalized head-related transfer functions (HRTFs) are essential for spatial audio rendering, but densely measuring an individual's HRTFs is costly and time-consuming. HRTF upsampling reduces this burden by estimating dense HRTFs from sparse measurements. Recent learning-based methods have achieved promising performance, but many remain tied to predefined measurement configurations. In this work, we propose GeoAtt, a variable-context HRTF upsampling framework that uses a single trained model across varying measurement configurations. GeoAtt performs geometry-aware, query-conditioned spatial aggregation over the available measurements independently at each frequency bin, followed by frequency-domain modeling using Conformer blocks. The relative geometry between the target and measured directions is incorporated as an additive bias in the cross-attention. Experiments on the SONICOM dataset show that a single trained model achieves the lowest log-spectral distortion across all four canonical Listener Acoustic Personalization (LAP) challenge measurement configurations and further generalizes to configurations that are not explicitly included during training.
著者のコメント
Submitted to IEEE ICASSP 2027
arXiv ID: 2609.25995 / 要約の誤りについて