ミームの判定を左右する小さな手掛かりを学ぶ分類法
Small Cues, Big Consequences: Learning Pivotal Cues for Multimodal Meme Classification
この論文をやさしく読む
ひとことで言うと
ミームの有害性や皮肉の判定で決め手になる、文字や画像の小さな手掛かりに注目する分類手法。
何に役立つ?
ミーム分類の評価データと分類器の改善に役立つ。実証結果はHarMeme、PrideMM、MemeCFでの比較実験である。
この研究の面白いところ
9,895件のミームに決定的な証拠の場所と理由を注釈し、単語と画像領域を必要な分だけ対応付ける構造を採用した。
どこまで分かった?
改善は記載された三つのデータセットと比較手法に対して報告されたもの。あらゆる文脈のミームで同じ判定精度が得られるとは要旨は示していない。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
ミームの有害性、憎悪表現、皮肉の意味は、画像、文字、あるいは両者の関係に含まれる小さいが決定的な手掛かりから生じることが多い。既存のマルチモーダル分類器は、画像と文字の全体的な表現に主に頼ると、このような証拠を見落としうる。本研究は、有害性、憎悪表現、皮肉にまたがる9,895件のミームからなる、手掛かりに焦点を当てたベンチマークMemeCFを導入する。各例には決め手となる証拠のモダリティと理由を注釈した。また、局所情報と全体情報を組み合わせるミーム分類器MemePIVOTを提案する。 MemePIVOTは、学習済みで固定したCLIPの特徴量を使い、非均衡最適輸送で単語と画像の領域を対応付ける。このとき無関係な証拠は対応付けずに残せる。さらに、証拠に基づく融合部によって、局所的な対応関係とミーム全体の文脈を不確かさを考慮しながら組み合わせる。HarMeme、PrideMM、MemeCFでの実験では、強力な文字のみ、画像のみ、マルチモーダル、視覚言語モデルの比較手法に対して一貫した改善を示した。データセットをまたぐ評価と構成要素を外した評価からも、決定的な証拠を明示的にモデル化することが頑健性を改善し、全体的なマルチモーダル表現を超える寄与を持つと分かった。コードとデータセットは公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-22(UTC)
- 最新改訂
- 2026-09-22 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-22 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Memes often derive their harmful, hateful, or sarcastic meaning from small but decisive visual, textual, or cross-modal cues. Existing multimodal classifiers can miss such evidence when relying mainly on global image-text representations. We introduce MemeCF, a cue-focused benchmark of 9,895 memes across harm, hate, and sarcasm, with annotations identifying the modality and rationale of the pivotal evidence. We also propose MemePIVOT, a local-global architecture for meme classification. MemePIVOT uses frozen CLIP features, unbalanced optimal transport to align words with image patches while allowing irrelevant evidence to remain unmatched, and an evidential fusion head to combine local grounding with global meme context under uncertainty. Experiments on HarMeme, PrideMM, and MemeCF show consistent gains over strong text-only, image-only, multimodal, and vision-language baselines. Cross-dataset and ablation results further show that explicit pivotal-evidence modeling improves robustness and contributes meaningfully beyond global multimodal representations. Our code and dataset are publicly available at https://github.com/AkshitSharma1/MemePIVOT
著者のコメント
Accepted to EMNLP 2026 Findings
arXiv ID: 2609.26907 / 要約の誤りについて