arXiv論文メモ
新着一覧
cs.LG · 査読状況未確認

異常データなしで異常検知モデルの性能を予測する

Can We Predict Anomaly Detection Performance from Embedding-Space Geometry?

Kevin Wilkinghoff and Zheng-Hua Tan

この論文をやさしく読む

ひとことで言うと

ラベル付き異常例がないとき、埋め込み空間の形から異常検知モデルを選ぶ方法を調べる。

何に役立つ?

異常例を集めにくい環境でモデルを比較する際に役立つ可能性がある。

この研究の面白いところ

正常例の分散だけでは不十分で、疑似異常を比較の基準に使うと選択が改善した。

どこまで分かった?

AUCの下界は正常例と異常例のスコア分離も含む理論式であり、正常データだけから直ちに真のAUCが分かるわけではない。実験は記載されたDCASE設定に基づく。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

異常検知システムは正常データだけで学習することが多い一方、モデル選択や評価には通常、ラベル付きの異常データが必要である。本研究は異常データを使わずに検知性能を予測できるかを調べる。k近傍法に基づく検知器について、受信者動作特性曲線下面積(AUC)の下界を導き、正常例と異常例のスコアの分離、およびそれぞれの分散と検知性能を結び付ける。局所スケーリングモデルの下でこの下界を用い、密度の変動、内在次元の不均一性、ドメイン間の不一致がスコア変動にどう寄与するかを特徴付ける。次に異常データを使わないモデル選択を調べ、正常例のスコア分散だけでは、異なる表現間の性能を確実には予測できないことを示す。この制限に対し、相対的なスコア分離を推定する基準を与える単純な疑似異常プローブを導入する。四つの埋め込みモデルと208個の候補システムを含むDCASE 2022~2025ベンチマークの実験では、疑似異常に基づく推定量が、異常データを使わないモデル選択を大幅に改善した。特に多様な疑似異常を使うと、ドメインが変わる条件で、従来の開発用データによる選択を上回った。埋め込み空間の幾何には異常検知性能を予測する情報がある一方、正常例だけに基づく性能推定は表現に依存することを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Anomaly detection systems are often trained using normal data alone, while model selection and evaluation typically require labeled anomalies. We study whether anomaly detection performance can be predicted without access to anomalous data. For kNN-based detectors, we derive a lower bound on the area under the ROC curve (AUC) that relates detection performance to the separation between inlier and outlier scores and to their respective variances. Under a local scaling model, we use this bound to characterize how density variation, intrinsic-dimensional heterogeneity, and cross-domain mismatch contribute to score variability. We then investigate anomaly-free model selection and show that inlier score variance alone does not reliably predict performance across different representations. To address this limitation, we introduce simple pseudo-anomaly probes that provide a reference for estimating relative score separation. Experiments on the DCASE 2022-2025 benchmarks, spanning four embedding models and 208 candidate systems, show that pseudo-anomaly-based estimators substantially improve anomaly-free model selection. In particular, diverse pseudo-anomalies enable anomaly-free model selection to outperform conventional development-set selection under domain shift. These results show that embedding-space geometry contains predictive information about anomaly detection performance while also highlighting the representation-dependent nature of inlier-only performance estimates.

arXiv ID: 2609.26460 / 要約の誤りについて