実録音でも動く音響合成器のパラメータ逆推定
Off-manifold robustness in synthesizer inversion with joint distribution flow matching
この論文をやさしく読む
ひとことで言うと
録音した音から音響合成器の設定を推定する際、正解設定のない実録音も学習に使える方法です。
何に役立つ?
実録音を使って音響合成器の設定を探す作業や、合成音と実録音の分布の違いへの対処に役立つと考えられます。
この研究の面白いところ
対応するパラメータのない実録音では音だけの分布を学び、対応付きの合成音では同時分布を学びます。推定時に参照音へ少し雑音を加えると再構成が改善しました。
どこまで分かった?
要旨での評価対象はSurge XTとDexedです。他の合成器や録音条件への適用、改善幅の数値は要旨に記載されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音響合成器の出力音から設定パラメータを逆に推定する研究では、音とパラメータの対応にある曖昧さを明示的にモデル化する生成モデルが、決定論的な手法を上回ると示されている。ただし、その学習には音とパラメータの組が必要で、通常は、抽出した設定や既定の設定を合成器に入力して音を生成する。このため、学習時と評価時に分布の不一致が生じ、正解パラメータの注釈が存在しない、合成器が作る音の分布から外れた実世界の録音では性能が低下し得る。 この問題を避けるため、著者らは、音とパラメータの同時分布を、モダリティごとに独立した雑音スケジュールを持つマルチモーダルな連続正規化フローでモデル化する。この定式化では、合成器で作った対応付け済みデータから同時密度と条件付き密度を学習できる一方、対応するパラメータのない実録音からは音の周辺分布だけを学習できる。これにより、パラメータのラベルなしで分布外の音をモデルに見せられる。さらに、モデルがあらゆる雑音水準で音からパラメータへの対応を学ぶため、推定時に参照音へ部分的に雑音を加えると、実録音の再構成が改善した。これは、大まかな構造を保ちながら、特定分布だけに見られる細部への感度を下げるという説明と整合する。Surge XTとDexedでの評価では、同時分布をモデル化すると、実世界の音と学習分布内の音の両方で逆推定が大きく改善した。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent work on synthesizer inversion shows that generative models outperform deterministic approaches by explicitly modeling the ambiguity in mapping audio to parameters. Training such models, however, requires audio-parameter pairs, which are typically obtained by rendering sampled or preset parameters through the synthesizer itself. This creates a train-test mismatch that can degrade performance on off-manifold real-world recordings, for which ground-truth parameter annotations do not exist. To circumvent this obstacle, we propose to model the joint distribution of audio and parameters with a multi-modal continuous normalizing flow using independent noise schedules for each modality. This formulation allows us to train joint and conditional densities with paired synthesizer data, while unpaired real recordings can train the audio marginal alone, exposing the model to off-manifold signals without requiring parameter labels. Further, because the model learns to map from audio to parameters at all noise levels, we find that partially noising the audio reference at inference improves real-audio reconstruction, consistent with reducing sensitivity to distribution-specific detail while preserving coarse structure. Evaluating on Surge XT and Dexed, we find that modelling the joint distribution substantially improves both inversion of real-world and in-domain audio.
arXiv ID: 2609.29320 / 要約の誤りについて