arXiv論文メモ
新着一覧
eess.AS · 査読状況未確認

音源の成分ごとに調整できるリアルタイム音楽修復

RESTORE: REal-time Steerable Music resTORation and bandwidth Extension via stem disentanglement

Meiying Chen, Benjamin R. Thompson and Michael C. Heilemann

この論文をやさしく読む

ひとことで言うと

古い録音のノイズや音楽成分を分け、利用者が調整しながら音を修復する方法。

何に役立つ?

歴史的な録音などで、除去する雑音と残す音を利用者が確認しながら決める作業に役立つ。

この研究の面白いところ

歌声やヒスノイズなどを一度の処理で分離し、高周波の補完で生成した音も独立に確認できる。

どこまで分かった?

要旨が報告する音質評価は多様な歴史的録音に基づく。提示された指標と速度以外の運用条件は要旨に記載されていない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

音声を修復するニューラル手法は通常、劣化した入力を単一のきれいな出力に変換する固定的なものとして設計される。その場合、どの音を除くかという判断が自動的に押し付けられ、修復後の信号に望まない音が加わることもある。何を修復済みの音とみなすかは主観的であるため、本研究は、音声修復を6種類の意味的な音源への分解として定式化し、利用者が処理をリアルタイムで対話的に制御できる枠組みRESTOREを導入する。事前学習済みのHTDemucsを拡張し、1回の順伝播で劣化した混合音を歌声、音楽、広帯域のヒスノイズ、衝撃的な過渡音、モデル化されていない残差に分離しながら、高周波帯域の拡張も同時に合成する。利用者は各音源成分の音量を調整して修復を制御でき、生成された音をほかの成分から分離して監査可能に保てる。RESTOREは多様な歴史的録音で比較手法より音質を改善し、Fréchet Audio DistanceはVGGishで12.13、CLAPで0.92まで低下した。また、美的な調整可能性についてスピアマンの順位相関係数が0.91を超え、単一GPUでリアルタイムの50倍の速度で動作した。コードと音声サンプルは論文が示すデモサイトで公開されている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-23(UTC)
最新改訂
2026-09-23 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Neural methods for audio restoration are typically framed as rigid mappings from degraded inputs to single clean outputs, enforcing decisions about what audio content is removed, and potentially adding unwanted content to the restored signal. Because what constitutes a restored audio signal is subjective, we introduce RESTORE, a framework that formulates audio restoration as a six-source semantic decomposition to allow for real-time interactive user control over the process. By expanding a pretrained HTDemucs backbone, a single forward pass disentangles a degraded mixture into vocals, music, broadband hiss, impulsive transients, and an unmodeled residual, while jointly synthesizing a high-frequency extension. Users may steer the restoration by adjusting stem gains, ensuring generative content remains isolated and auditable. RESTORE improves audio quality on diverse historical recordings compared to baselines,lowering Frechet Audio Distance (FAD) (12.13 VGGish; 0.92 CLAP) and delivering aesthetic steerability (Spearman rho greather than 0.91) at 50x real-time on a single GPU. Code and audio samples are available at https://melissachen15.github.io/restore-audio-demo.

著者のコメント

Code and audio samples: https://melissachen15.github.io/restore-audio-demo

arXiv ID: 2609.28683 / 要約の誤りについて