arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

背景知識から段階的に学び、差別的ミームの判断を説明する

Learn Before You Judge: Progressive Knowledge-to-Decision Alignment for Explainable Hateful Meme Detection

Bo Xu, Chenyuan Wang, Xinyu Chen, Quanhao Zhu, Rui Lin, Liang Zhao, Hongfei Lin, Feng Xia

この論文をやさしく読む

ひとことで言うと

画像と文章を組み合わせたミームの有害性を、背景知識の学習から段階的に判断できるようにする方法です。

何に役立つ?

理由と根拠を伴うコンテンツ判定の設計に役立ちます。背景を理解する学習とラベルを当てる学習の干渉を減らす狙いがあります。

この研究の面白いところ

背景知識、憎悪表現の検出、判定境界の調整を3段階で学び、各段階では一つの目的に集中します。説明生成と検出を一度に最適化する方式との違いが中心です。

どこまで分かった?

公開3ベンチマークで最先端の検出成績と根拠付き説明を報告しています。要旨には具体的なスコアや言語・文化圏をまたぐ結果はなく、あらゆる文脈の正しい判断を保証するものではありません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

差別的なミームは、画像と文章の暗黙的な相互作用を通して攻撃的な内容を広め、オンラインコミュニティーの安全に深刻な脅威をもたらす。近年、マルチモーダル大規模言語モデルが差別的ミームの検出に広く使われ、説明可能な検出結果の生成にも採用されるようになっている。しかし、既存の「説明してから検出する」方法は、説明生成とラベル予測を同じ学習過程で結びつけることが多いと分かった。この結びつきは課題の目的間に干渉を生じさせ、検出性能を制限し、単純な教師あり微調整(SFT)の基準手法より悪い結果を招くことさえある。 この課題に対し、説明可能な差別的ミーム検出のための、知識から判断への段階的整合手法ProKDAを提案する。人間のアノテーション担当者の訓練過程に着想を得て、まずエージェント型の背景知識構築処理系を用い、ミームの理解に関係する外部知識を得る。その後、背景知識の学習、差別性検出の学習、差別性の判定境界の整合を順に行う3段階の学習戦略を採る。 両課題を共同最適化する従来手法とは異なり、ProKDAは各段階で単一の学習目的に集中する。この設計によって2つの課題間の干渉を減らし、背景知識を頑健な検出判断へと段階的に変換する。3つの公開差別的ミーム評価ベンチマークでの実験により、ProKDAは最先端の検出性能を達成し、差別的ミームのモデレーションに向けて、正確で説明可能かつ根拠に支えられた判断を提供することが示された。プロジェクトページ:https://meizhiyuan88666.github.io/prokda。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explainable detection results. However, we find that existing explain-then-detect methods often couple explanation generation and label prediction within the same training process. This coupling causes interference between task objectives, leading to limited detection performance and even worse results than simple SFT baselines. To address these challenges, we propose ProKDA, a progressive knowledge-to-decision alignment method for explainable hateful meme detection. Inspired by the human annotation training process, ProKDA first uses an agentic background knowledge construction pipeline to obtain external knowledge related to meme understanding. It then adopts a three-stage training strategy that sequentially performs background knowledge learning, hatefulness detection learning, and hatefulness boundary alignment. Unlike prior explain-then-detect methods that jointly optimize both tasks, ProKDA focuses on a single training objective at each stage. This design reduces interference between the two tasks and progressively transforms background knowledge into robust detection decisions. Experiments on three public hateful meme benchmarks show that ProKDA achieves state-of-the-art detection performance and provides accurate, explainable, and evidence-supported decisions for hateful meme moderation. Project page: https://meizhiyuan88666.github.io/prokda.

著者のコメント

26 pages, 16 figures, 7 tables

arXiv ID: 2609.19778 / 要約の誤りについて