arXiv論文メモ
新着一覧
cs.CV / cs.MM · 掲載先の記載あり

説明キーワードを使って有害ミームを分類する

MemeTAG: Keyword-Driven Meme Classification through Tag Embedding Reconstruction

Akshit Sharma, Prashant W. Patil

この論文をやさしく読む

ひとことで言うと

画像と文章の組み合わせが持つ意味を説明キーワードにまとめ、その情報を使ってミームを分類します。

何に役立つ?

有害ミームの分類支援への利用が考えられます。報告された成果は三つのデータセットでの分類性能であり、実運用のモデレーション効果とは区別が必要です。

この研究の面白いところ

生成したキーワードをそのまま分類器に渡すだけでなく、一つの意味表現に集約し、その再構成を補助学習目標にしています。

どこまで分かった?

要旨には性能の具体的な数値や誤分類の内訳、対象言語・文化をまたぐ評価は示されていません。最高性能という記述は著者による当該データセットでの比較結果です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

有害なインターネットミームの拡散は重大な社会的脅威となっているが、その内容は微妙な意味合いを持ち、複数の情報様式にまたがるため、自動分類は依然として大きなアルゴリズム上の課題である。この課題に対し、キーワードを考慮したミーム分類手法を切り開く、二つの目的を持つ新しいフレームワークMemeTAGを提案する。 中心となる工夫は、二段構成の意味的な誘導機構である。まず、事前学習済みの視覚言語モデルを利用し、高次の意味を捉えた説明キーワードの集合を生成する。次に、注意機構に基づくAggregated Tag Inference Network(ATIN)を導入し、これらのキーワードを、豊かな意味を持つ単一の埋め込み表現へと集約する。この埋め込みを新しい補助的な再構成損失の目標とし、視覚特徴とテキスト特徴が深く整合するようモデルに学習させる。この手法を効率的な三段階の学習戦略と組み合わせることで、HarMeme、Hateful Memes Challenge(HMC)、PrideMMの各データセットで新たな最高水準の性能を達成し、既存の最先端手法を明確に上回る。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-17(UTC)
最新改訂
2026-09-17 · v1
査読・掲載
掲載先の記載あり

著者による掲載先の記載:Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026, pp. 7679-7688。出版社での独立確認は未実施です。

arXivで読むPDFDOI

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

The proliferation of harmful internet memes poses a significant societal threat, yet their automated classification remains a formidable algorithmic challenge due to the nuanced, multimodal nature of their content. To address this, we introduce MemeTAG, a novel dual-objective framework that pioneers a keyword-aware approach to meme classification. Our core innovation is a two-part semantic guidance mechanism: first, we leverage a pretrained Vision-Language Model to generate a set of descriptive keywords, that capture the high-level semantics. Second, we introduce the Aggregated Tag Inference Network (ATIN), an attention-based module that distills these keywords into a single, rich semantic embedding. This embedding serves as a target for a novel auxiliary reconstruction loss, which compels the model to learn deeply aligned visual and textual features. This approach, combined with an efficient three-stage training strategy, establishes a new state-of-the-art on the HarMeme, Hateful Memes Challenge (HMC), and PrideMM datasets, decisively outperforming existing state-of-the-art methods.

著者のコメント

10 pages, 3 figures; published in the Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

arXiv ID: 2609.20962 / 要約の誤りについて