有害ミーム判定の失敗を内部表現と出力経路に分ける
Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection
この論文をやさしく読む
ひとことで言うと
有害な画像投稿をAIが見誤るとき、内部に情報がないのか、情報はあっても答えに使えていないのかを分けて調べます。
何に役立つ?
有害コンテンツ判定モデルの弱点を診断し、読み出しや較正を改善する手掛かりになります。内部情報の存在と通常の判定能力を分けて評価できます。
この研究の面白いところ
内部特徴を読むだけでなく、特徴の除去・置き換えや出力回復を試しています。画像と文章の組合せに由来する信号であることも対照実験で調べています。
どこまで分かった?
教師ありプローブの性能向上は、モデルがもともと正しい判断規則を持っていた証拠ではありません。感度の倍率は採用スコア尺度に依存し、複数課題をまとめた適応では負の転移も報告されています。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模視覚言語モデルが有害なミームを誤分類するとき、その原因は、内部に必要な証拠がないことか、表現されている証拠を出力へ導けないことかもしれない。本研究ではGemma-3とQwen3.5について、疎自己符号化器、役割条件付きプローブ、因果介入、回復実験を使い、6つの有害コンテンツベンチマークでこの2つを区別する。さらにスペイン語と、ヒンディー語・英語を混ぜた評価も行う。 疎な読み出しは主要な2値分類課題6つすべてでモデル本来の予測を上回った。Qwenの平均マクロF1は、本来の0.432に対して0.740となり、残差再構成では0.486に達した。Gemmaは0.532から0.714へ改善した。これらの差は、教師ありの読み出しで情報にアクセスできることを示し、既存の本来的な判定規則があることを示すものではない。また、最も影響力のあるトークンの役割は課題によって異なる。 評価したスコア尺度では、Qwenの出力に現れない特徴の除去はプローブに対して24〜63倍敏感である一方、文字どおりyes/noを答える課題での出力へつながる特徴のパッチングは、出力に対して16〜140倍敏感である。較正のみを用いた経路調整は平均差の93.3%を回復し、プローブから蒸留したLoRAはモデル本来の予測を改善するが、共有のマルチタスク適応は負の転移を起こす。 Facebook Hateful Memes上のGemma-3-12Bの事例研究では、分散したランク32の画像・プロンプト相互作用が見つかり、マクロF1は本来の0.685に対して0.756となった。頑健性の対照実験は、この信号が英語以外にも及び、付随するOCRだけでは説明できず、組み合わされた視覚的証拠に依存することを示す。したがって、有害ミームの分類では、表現だけでなく、出力への経路が繰り返し現れる障害となっている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-16(UTC)
- 最新改訂
- 2026-09-16 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-16 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its output. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages $0.740$ versus $0.432$ native macro-F1, while residual reconstruction reaches $0.486$, whereas Gemma improves from $0.532$ to $0.714$. These differences reflect supervised accessibility rather than a pre-existing, native decision rule, and the most influential token role depends on the task. Under the evaluated score scales, Qwen silent-feature ablation is $24-63$ times more probe-sensitive, whereas routed-feature patching on literal yes/no tasks is $16-140$ times more output-sensitive. Calibration-only routing recovers $93.3$% of the mean gap, and probe-distilled LoRA improves native predictions, although shared multi-task adaptation causes negative transfer. A case study of Gemma-3-12B on Facebook Hateful Memes finds a distributed rank-32 image-prompt interaction, reaching $0.756$ versus $0.685$ native macro-F1. Robustness controls show that the signal extends beyond English, is not explained solely by accompanying OCR, and depends on paired visual evidence. Thus, routing, rather than representation alone, is a recurring bottleneck in harmful meme classification.
著者のコメント
40 pages, 9 figures
arXiv ID: 2609.18860 / 要約の誤りについて