LLMによる国際比較調査で文化差の縮小と誇張を測る
Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations
この論文をやさしく読む
ひとことで言うと
LLMで国ごとのアンケート回答を模擬するとき、文化差が薄まったり誇張されたりしていないかを測る指標です。
何に役立つ?
合成回答を国際比較に使う前の評価に役立つと考えられます。要旨では従来の分布指標が見逃す国間の差をCDPが捉えたと報告しています。
この研究の面白いところ
各国の回答分布がよく合って見えても、国同士の違いは平板化している場合がありました。
どこまで分かった?
評価は四モデル、三プロンプト法、二調査分野で行われました。CDPには一度の人間回答による較正が必要です。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)は、集団の回答分布を推定するため、合成された調査回答者として使われることが増えている。文化をまたぐ調査のシミュレーションでは、各国内の分布を忠実に再現するかだけでなく、国と国の違いが保たれるかも評価すべきである。しかしJensen–Shannonダイバージェンス(JSD)のような既存の距離指標は、国間の違いを直接捉えない。この問題に対処するため、一度だけ人間の回答で較正する、参照情報の少ない診断指標Cultural Divergence Preservation(CDP)を導入する。CDPは国間の差が小さくなることを文化的な平板化、大きくなることを文化の戯画化として捉える。 四つのLLM基盤モデル、三つの人物像に基づくプロンプト法、世界価値観調査(WVS)とビッグファイブ性格検査の二つの調査分野で実験した。結果は、従来の分布忠実度指標とCDPの間に体系的な食い違いがあることを示した。制御実験では、国間の差を弱めたり強めたりするとCDPは単調に変化したが、対応するJSDの変化は比較的小さかった。実際のLLM生成の監査では、DeepPersona-Inspiredというプロンプト法は従来指標に好まれることが多い一方、すべてのモデルと分野の組み合わせで最も強い平板化を示した。CDPは国間の差の縮小や拡大を直接定量化し、従来の忠実度指標を補う。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-24(UTC)
- 最新改訂
- 2026-09-24 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-24 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration. CDP identifies reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature. To evaluate CDP, we conduct experiments across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test. The results reveal a systematic discrepancy between conventional fidelity metrics and CDP. Controlled experiments show that CDP changes monotonically as cross-country divergence is attenuated or amplified, while the corresponding changes in JSD remain relatively small. In our audit of real LLM generations, DeepPersona-Inspired prompting is frequently favored by conventional fidelity metrics but exhibits the strongest flattening in every model--domain block. CDP thus complements fidelity metrics by directly quantifying the attenuation or amplification of cross-country divergence.
著者のコメント
Accepted to the EMNLP 2026 Workshop on Pluralistic AI & NLP (PANDORA)
arXiv ID: 2609.29928 / 要約の誤りについて