人の演奏とAI生成音源が混ざる楽曲を検出する
A Stem-Agnostic Approach to Hybrid AI Music Detection
この論文をやさしく読む
ひとことで言うと
人の演奏とAI生成の音源が混ざった楽曲で、どのパートが生成されたか調べる方法を提案した。
何に役立つ?
混合楽曲のAI生成部分を調べる際の技術的な手掛かりになる。
この研究の面白いところ
音の時間・周波数の各位置で合成らしさを示す表現を作り、一つのモデルで複数パートを扱った。
どこまで分かった?
ボーカルなどでは良好だったが、ベースでは難しく、音源分離の品質が検出の制約となる。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
音楽制作で生成音声が使われるようになり、人間の演奏とAI生成の音源パートを混ぜた楽曲が増えている。これは楽曲全体を二値で判定する従来のAI音楽検出器にとって難しい。本研究は、混合された楽曲の中にある合成音源を、音源パートの種類を限定せず識別する枠組みを提案する。音声スペクトル上に合成内容の局所的な確率を写し出す、新しい時間・周波数表現「inspectrogram」を導入する。これと、対象音源パートのエネルギー優位度を推定するWienerフィルターを組み合わせ、一つの畳み込みニューラルネットワークで特定のパートが生成されたものかを評価する。 合成した混合楽曲で学習し、複数の音源パートの種類で評価した。ボーカル、ドラム、ギターのような高周波の音源では良好な性能を得た一方、低周波で帯域の狭いベースでは苦戦した。音源分離の質が検出精度に影響すると結論付け、音源分離を主なボトルネックであり、今後の重要な研究方向だとする。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The inclusion of generative audio in the music production process has led to an increase in hybrid music tracks that blend authentic human performances with AI-generated stems, challenging traditional AI music detectors which operate in a binary setting. In this work, we propose a stem-agnostic framework for identifying synthetic audio sources within hybrid musical mixtures. We introduce the inspectrogram, a novel time-frequency representation that maps localized probabilities of synthetic content across the audio spectrum. By combining the inspectrogram with a Wiener filter estimating target stem energy dominance, a single CNN model evaluates whether the specific stem is generated. Trained on rendered hybrid mixtures and evaluated across various stem classes, our model achieves strong performance on high-frequency sources such as vocals, drums, and guitar, but struggles on the low-frequency, narrow-band bass. We conclude that the quality of separation impacts the detection accuracy and identify source separation as a primary bottleneck and a crucial direction for future research.
著者のコメント
4 pages + references, 5 figures, submitted to the 2027 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)
arXiv ID: 2609.26956 / 要約の誤りについて