楽譜を手掛かりに混合音から音符ごとの波形を分離
On a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership
この論文をやさしく読む
ひとことで言うと
楽譜の音高とタイミングを手掛かりに、混合録音から個々の音符の波形を取り出す方法を示した。
何に役立つ?
録音の音符単位の分析や編集に使える可能性がある。
この研究の面白いところ
音符ごとの抽出後に、同時に鳴る音符間で混合音のエネルギーを再配分する二段階の構成。
どこまで分かった?
評価は整備したSCNS-Evalで行い、楽譜とライブラリは学習用と分離した。実際の多様な商用録音での性能は要旨にない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
楽譜を手掛かりにする音符分離は、多声音源の録音などから、演奏された一つひとつの音符の波形を取り出すことを目指す。既存の深層学習システムは一般に楽器単位の音源分離だけを対象としている。著者らは、知る限り初めての、楽譜情報に基づく音符分離の深層学習手法NoteSepを提案する。 NoteSepは、音符ごとに抽出モデルNoteGrabを一度ずつ適用して指定した音符を取り出す。音高、開始時刻、終了時刻を条件に、双方向の交差注意で結ばれた2つのU-Netが倍音成分と打楽器的成分を分離する。選択的な倍音ゲートは、打撃の立ち上がりを保ちながら低いオクターブからの干渉を抑える。最後に、Adaptive Set Ownershipを用いる共同分離段階で、同時に鳴る音符のNoteGrab推定を比較し、混合音のエネルギーを再配分する。 学習用に混合音25,729件と対象音743,920件からなるSCNS-Trainを作り、評価用には16楽器を含み、楽譜と音源ライブラリが学習用と重ならないSCNS-Evalを整備した。SCNS-EvalでNoteSepのSI-SDR中央値は7.39 dBで、最も強い比較手法の2.49 dBを上回った。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39~dB, compared with 2.49~dB for our strongest baseline. See the demo page at https://benschou.com/notesep.
著者のコメント
Under Review
arXiv ID: 2609.29071 / 要約の誤りについて