中国のAI生成コンテンツ規制に対する言語モデルの評価
Benchmarking LLM Compliance with China AI Generated Content Regulations
この論文をやさしく読む
ひとことで言うと
中国のAI生成内容に関する要件を著者らが評価項目にし、中国語質問で20の言語モデルを比較しています。
何に役立つ?
法規制を意識した出力の評価方法や、拒否率と適合率の違いを考える資料になります。6分野、2,303問を使っています。
この研究の面白いところ
203問の独自の憲法関連質問を含み、複数の判定者が独立に評価します。海外モデルも高い適合傾向を示し、差が大きいのは思想的整合に近い分野だと報告しています。
どこまで分かった?
これは論文の評価枠組みと結果の紹介です。実際の法的適合を独立に認定したものではなく、要旨には判定者の誤り率や具体的なモデル別数値は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)の普及に伴い、コンテンツの規制適合性に関するリスクが高まっている。先行研究は英語の文脈でこのリスクへの対処に貢献してきた一方、中国語コンテンツの複雑さを十分に重視してこなかった。本論文は、中国の現行のAI生成コンテンツに関する適合要件に沿って、主要な20のLLMの評価結果を示し、中国の規制状況に関する知見を提供する。 6つの異なる側面にわたる2,303問を用い、適合率と回答拒否率を評価する新しい枠組みを設計する。この中には、独自に作成した憲法に関する203問が含まれる。枠組みは複数の判定役を用い、それぞれの階層的な整合性メモリに基づいて独立に判定を生成させる。結果は、標準的な中国語の質問を用いていても、海外のモデルが高い適合性を示すこと、主な違いはイデオロギー上の整合性と密接に関係する側面に由来する可能性があることを示す。 本研究は、法的根拠を持つ共通の適合要件の下で、中国のLLMとそれ以外のLLMの双方を世界のAIコミュニティが評価できる、規制に関するベンチマークを構築する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
The widespread adoption of LLMs has led to escalating content compliance risks. Prior works have contributed to addressing these risks in the English context, downplaying the complexity of Chinese language content. This paper follows China's current AI-Generated content compliance requirements and provides evaluation results on 20 notable LLMs, offering insight into China's regulatory landscape. We design a novel framework to assess the compliance and refusal rates with 2303 questions spanning six distinct dimensions, including 203 self-constructed constitutional questions. The framework employs several judges to generate verdicts independently based on their hierarchical alignment memory. Our findings show that international models also exhibit high levels of compliance despite the use of standard Chinese questions, and the main differences may stem from dimensions closely related to ideological alignment. We establish a regulatory benchmark that enables the global AI community to evaluate both Chinese and non-Chinese LLMs under a unified set of legally grounded compliance requirements.
著者のコメント
5 pages, 3 figures, with appendix still improving
arXiv ID: 2609.19989 / 要約の誤りについて