arXiv論文メモ
新着一覧
eess.AS / cs.SD / eess.SP · 査読状況未確認

少数の実録音で話者距離推定を補正する方法

Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

Michael Neri and Archontis Politis and Tuomas Virtanen

この論文をやさしく読む

ひとことで言うと

合成音で学習した話者距離推定を、少数の実録音を使って補正する研究です。

何に役立つ?

実録音の距離ラベルが少ない場合に、モデルを再学習せず距離推定を調整する手がかりになります。

この研究の面白いところ

推定値の絶対的な誤差より、近い・遠いの順序を保てるかが較正には重要だと分析します。

どこまで分かった?

評価は三つの実録音コーパスです。要旨には必要なラベル数や改善率の具体値は示されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

話者とマイクの実際の距離を付記した録音は少ないため、話者距離の推定器はほぼすべて模擬した室内音響で学習される。本研究は、そのように学習したモデルが実録音へうまく移らないことを示す。評価した三つの実録音コーパスでは、コーパスの平均距離を常に予測するだけの方法が、学習したどのモデルよりも正確だった。そこで、合成音で学習し固定した推定器を役立つものにするには、距離ラベル付きの実録音が何件必要かを問い、勾配計算や再学習なしに出力の尺度を直す事後的な較正関数を調べる。達成可能な誤差の解析によると、較正の成否を制限するのは推定器の絶対精度ではなく、発話を距離順にどれほどうまく並べられるかである。一定の偏りや誤った出力尺度は較正そのもので正確に除去できるためだ。少数の標本から各係数を推定する費用との釣り合いを取り、コーパスとラベル付けの予算に応じてどの較正関数が有利かを説明する基準と、明確な選択を必要としない縮小推定の変種を得る。結果は、合成音で学習したモデルの保存時点を選ぶ際、絶対誤差より真の距離との線形相関を用いることを示唆する。コード、データセット、解析も公開している。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.

著者のコメント

Submitted to IEEE International Conference of Acoustics, Speech, and Signal Processing (IEEE ICASSP 2027)

arXiv ID: 2609.29203 / 要約の誤りについて