画像・文章の学習データに混入した汚染例を特徴の順位から除去
TraceGuard: Adaptive Multimodal Poison Filtering through Cross-Feature Rank Agreement
この論文をやさしく読む
ひとことで言うと
画像と文章の学習データに紛れた、モデルの振る舞いを誘導する汚染例を検出・除去する方法です。
何に役立つ?
外部データでマルチモーダルモデルを学習する前のデータ点検に役立つと考えられます。19の攻撃設定で平均98.4%の汚染例を除去しました。
この研究の面白いところ
個々の画像・文章の不自然さだけでなく、集合全体で繰り返すパターンと複数特徴の順位の一致を利用します。
どこまで分かった?
清浄な例も平均5.4%除去され、適応的攻撃では検出失敗、汚染のない集合では不要な除去が報告されています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
マルチモーダルモデルの学習は外部から集めた画像・文章のデータに依存しており、攻撃者がデータを汚染する機会がある。巧妙な攻撃ではもっともらしい画像と文章の組を保ちながら、検出器が使う差異を隠せるため、一見きれいなデータでも学習済みモデルを誘導できる。そこで本研究は、攻撃が有効であり続けるために汚染集合が何を保つ必要があるかを問う。少数の汚染例でも、学習中に攻撃者の狙う振る舞いを引き起こすだけの集団的な影響力が必要である。この影響を、攻撃パターンの出現頻度と、それを含む例が共同でモデルに与える作用の強さから分析する。この分析から、被害モデルを学習せずに、画像と文章をまたぐ近傍、繰り返される文言、文章の一部を消去した後の変化を調べる、データ集合全体の6つの特徴を導く。TraceGuardは、相補的な特徴による順位の一致を使って疑わしい例を見つける適応的な順位ベースのフィルタである。共通パターンを使って選択集合を洗練し、攻撃内容や汚染率を知らなくても、データ集合ごとに除去の閾値を調整する。画像・文章学習、生成型視覚言語モデルの微調整、エンコーダーの転移テストを含む19の攻撃設定で、汚染例の平均98.4%と清浄な例の5.4%を除去した。フィルタ後のデータで学習すると、13設定では残る攻撃指標が最大1%だった。除去数を合わせた対照実験と要素除去実験は、例の選択と適応的な除去の寄与を支持した。一方、負荷試験では、適応的な攻撃での検出失敗と、汚染のないデータでの不要な除去も確認された。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We therefore ask which properties a poison set must preserve for the attack to remain effective. A small poison set must still exert enough collective influence during training to induce the attacker's target behavior. We analyze this influence in terms of how often an attack pattern occurs and how strongly the examples carrying it jointly affect the model. This analysis motivates six corpus-level features that examine cross-modal neighborhoods, recurring text, and changes after text-span erasure without training the victim model. We introduce TraceGuard, an adaptive rank-based filtering method that uses agreement among complementary feature rankings to identify suspicious examples. It refines the selected set through shared patterns and adapts the removal threshold to each corpus without knowing the attack or poison rate. Across 19 attack configurations spanning image-text learning, generative vision-language model fine-tuning, and encoder-transfer tests, TraceGuard removes an average of 98.4% of poisoned examples and 5.4% of clean examples. After training on the filtered corpora, the residual attack metric is at most 1% in 13 configurations. Matched-removal controls and ablations support the contributions of sample selection and adaptive removal. Stress tests also identify detection failures under adaptive attacks and unnecessary removal on poison-free corpora.
著者のコメント
42 pages
arXiv ID: 2609.29099 / 要約の誤りについて