少数の検証済みデータから学習データの汚染を見つける
SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples
この論文をやさしく読む
ひとことで言うと
専門家が確認した少数の正常例と汚染例を手掛かりに、学習データへ混入した攻撃用データを探す防御手法です。別のデータで学んだ特徴の類似度を使います。
何に役立つ?
大量のデータを人手で検査する負担を減らし、学習前のデータ点検を支援する用途が考えられます。7種類の攻撃を使ったベンチマークでは、少数でも汚染例の確認が有効でした。
この研究の面白いところ
正常と確認した例だけでなく、汚染と確認した例も基準に使う点が特徴です。確認済み例を単純に増やすより、正常例を各クラスへどう配分するかが重要という結果も報告しています。
どこまで分かった?
少数であっても、専門家による正確な検証と別データで学習した特徴抽出器を必要とする設定です。要旨には検出率、誤検出率や具体的な検証件数は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
機械学習が公開された信頼できないデータ源への依存を強めるにつれ、選んだ標的の誤分類を引き起こすために悪意ある例を学習データへ混入するデータポイズニング攻撃の脅威が増している。既存の防御策は、どの例が汚染されているかという正解情報がまったくないことを仮定するか、汚染されていないと確認済みの大量の例を利用できることを仮定する。後者の仮定を満たすには大きな費用がかかる。信頼できる検証には、多くの資源や人手が必要となり得るからである。この費用は、汚染例を見た目で正常データと区別できないクリーンラベル攻撃で特に高い。 大量の検証済み例を要求するのは現実的ではないため、正常例と汚染例の両方を含む少数の検証済み例に依拠することを提案する。すなわち、各例はフォレンジックの専門家による調査を通じて、正常か汚染済みかのいずれかが確認されている。この場合の課題は、ほとんどの分類モデルが過学習してしまうほど小さな検証済み集合に基づき、汚染を検出することである。 この課題に対して、正解情報に基づく除外のための類似度ベースの手法SAGEを提案する。別のデータセットで汎用的な特徴抽出器を学習した後、検証済み集合に基づく、類似度で重み付けしたノンパラメトリックな予測によって、汚染された学習例に印を付ける。標準ベンチマーク上で7種類のクリーンラベル攻撃手法に対して検証し、検証済みの汚染例をほんの数件利用できるだけでも、大きな利点が得られることを示す。また、検証済みの正常例が各クラスにどのように分布しているかは、検証済み例の総数よりも重要であることが分かった。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-10-01(UTC)
- 最新改訂
- 2026-10-01 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-10-01 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.
arXiv ID: 2610.01788 / 要約の誤りについて