中国ネット流行語を英語で理解できるかを測る評価集
CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords
この論文をやさしく読む
ひとことで言うと
中国のネット流行語について、英語で意味を説明し、近い表現を選び、有害性を判断できるかを調べる評価集です。直訳だけでは分からない文化的な意味を対象にしています。
何に役立つ?
多言語モデルが俗語や婉曲表現をどこまで理解できるかを調べるために使えます。翻訳の意味理解と、有害な内容の検出を別々の課題として評価できます。
この研究の面白いところ
同じ流行語に意味説明、英語の対応表現、分類、有害性の注釈を付けています。単に正答できるかだけでなく、選択肢を変えても対応付けが安定するかを問います。
どこまで分かった?
対象は中国語から英語への3,001語の評価です。要旨にはモデルごとの得点はなく、他の言語やすべての新しい流行語での性能まで示したものではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
中国のソーシャルメディアでは、インターネット上の流行語が膨大かつ絶えず変化する語彙を形成している。その意味はしばしば字義通りではなく、地域固有の文化や語用論的な文脈に深く根差している。既存研究は主に中国語の中でこれらの流行語を解釈することに集中しており、大規模言語モデル(LLM)が文化に根差した知識を言語間で転移し、意図された意味を英語で正確に伝えられるかは、ほとんど検討されていない。この言語横断能力は安全性にとっても重要である。有害な表現は、文化固有の同音語、婉曲表現、皮肉、暗号的な言い回しを通じて、攻撃的な内容を覆い隠す場合があるためだ。 本論文では、高度なLLMが中国のネット流行語を言語横断的に理解する能力を調べる。そのために、中国のネット流行語を中国語から英語へ理解する能力を測る初のベンチマーク、CIBuzzBenchを導入する。CIBuzzBenchは3,001語の中国ネット流行語からなり、英語による意味説明、対応する英語表現、カテゴリラベル、有害性ラベルを付与している。この注釈に基づき、意味説明、言語横断的な対応表現の照合、文化的文脈に基づく有害性検出の三つの評価課題を設計する。 代表的な最先端のプロプライエタリLLMおよび中国のLLMを、英語と中国語の両方で指示する設定で評価する。結果は、LLMが中国のネット流行語の言語横断理解に依然として苦戦していることを示す。特に、細かな非字義的解釈、選択肢の変更に対して頑健な対応表現の照合、適切に較正された有害性検出が課題となっている。これらの知見は、文化に根差した言語現象が、多言語LLMと安全性を重視した評価に継続的な難しさをもたらすことを示している。データセットとコードはhttps://github.com/SuperYFan/CIBuzzBenchで公開されている。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzzwords within Chinese, leaving largely unexplored whether LLMs can transfer such culturally grounded knowledge across languages and accurately convey the intended meanings in English. This cross-lingual capability is also critical for safety, as harmful expressions may obscure their offensive content through culture-specific homophony, euphemism, irony, or coded language. In this paper, we investigate the ability of advanced LLMs to understand Chinese internet buzzwords across languages. To this end, we introduce CIBuzzBench, the first benchmark for cross-lingual Chinese-to-English understanding of Chinese internet buzzwords. CIBuzzBench comprises 3,001 Chinese internet buzzwords annotated with English meaning explanations, English equivalents, category labels, and harmfulness labels. Based on these annotations, we design three evaluation tasks: Meaning Explanation, Cross-lingual Equivalent Matching, and Culturally Grounded Harmfulness Detection. We evaluate representative state-of-the-art proprietary and Chinese LLMs under both English- and Chinese-prompting settings. Our results show that LLMs continue to struggle with the cross-lingual understanding of Chinese internet buzzwords, particularly in fine-grained non-literal interpretation, robust equivalent matching under option perturbations, and calibrated harmfulness detection. These findings highlight the persistent challenges posed by culturally grounded language phenomena for multilingual LLMs and safety-oriented evaluation. The dataset and code are available at https://github.com/SuperYFan/CIBuzzBench.
arXiv ID: 2609.21722 / 要約の誤りについて