arXiv論文メモ
新着一覧
cs.SD / eess.AS · 査読状況未確認

生成元やデータの来歴を考慮してAI音楽検出器を評価する

ArtifactBench: Lineage-Aware Evaluation of AI-Generated Music Detectors under Distribution Shift

Heewon Oh

この論文をやさしく読む

ひとことで言うと

AIが作った音楽を見分ける検出器を、音源の来歴や派生版の重複、推論できなかった曲まで区別して比べる評価基盤です。

何に役立つ?

検出器を比較するとき、見かけの高スコアがデータ重複や一部の曲だけでの評価に左右されていないかを検討するのに役立ちます。

この研究の面白いところ

分類の誤りと、そもそも推論ができなかったことを分けます。モデルが共通して処理できた音源だけで比較する条件や、閾値の選び方による順位変化も扱います。

どこまで分かった?

報告された0.982などの数値は、共通して推論に成功した562曲でのものです。すべての音楽や生成器で同じ性能が得られるという意味ではなく、要旨自体が大きな分布変化を報告しています。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

AI生成音楽の検出器は、訓練データとの重複、生成器の系統、音源の出所、音声変換の履歴が一部しか分からないベンチマークの集計スコアで比較されることが多い。本論文では、生成器の系列とバージョン、実際の音楽の領域、収集コホートの変化、推論の実行可能範囲にわたって検出器の挙動を測る、系統情報を考慮した評価基盤ArtifactBenchを導入する。 このベンチマークは、元の録音とその派生版を同じ内容としてグループ化し、校正と最終テストを分離し、推論失敗を分類誤りとは独立に記録する。また、全体の集計指標に加えて、元音源単位の性能を不確実性とともに報告する。複数の公開検出器を、バージョンを固定した共通の手順で評価し、情報漏洩の制御、コホートの利用可能性、閾値の方針、モデルごとの欠測が、測定性能とモデル順位をどう変えるかを検討する。 全モデルが共通して推論に成功したテスト音源562曲の集合では、ArtifactNetのAUROCは0.982、バランス正解率は0.918となり、公開Deezer検出器の0.761/0.776と比較される。SpecTTTraとCLAMは、この分布が変化したコホートではAUROCが0.30未満となった。これらの結果は、集計スコアだけでは隠れてしまう、生成器ごと、および実音楽の領域ごとの大きな分布変化も明らかにしている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-20(UTC)
最新改訂
2026-09-20 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

AI-generated music detectors are commonly compared using aggregate scores on benchmarks whose training overlap, generator lineage, source provenance, and audio-transformation history are only partially observable. This paper introduces ArtifactBench, a lineage-aware evaluation suite for measuring detector behavior across generator families and versions, real-music domains, collection-cohort shift, and inference coverage. The benchmark groups source recordings and their derived variants by content identity, separates calibration from final testing, records inference failures independently from classification errors, and reports source-level performance with uncertainty in addition to aggregate metrics. We evaluate multiple publicly available detectors under a version-pinned common protocol and examine how leakage control, cohort availability, threshold policy, and model-specific missingness alter measured performance and model ranking. On the 562-track common-success test intersection, ArtifactNet obtains 0.982 AUROC and 0.918 balanced accuracy, compared with 0.761/0.776 for the public Deezer detector; SpecTTTra and CLAM fall below 0.30 AUROC under this shifted cohort. These results also expose substantial generator- and real-domain shifts that aggregate scores alone conceal.

著者のコメント

9 pages, 1 figure, 6 tables. Related to ArtifactNet (arXiv:2604.16254); this work contributes the benchmark and evaluation protocol rather than a new detector

arXiv ID: 2609.23550 / 要約の誤りについて