arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

LLMによる国際比較調査で文化差の縮小と誇張を測る

Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations

Yeeun Chae, Yewon Choi, Seunghyun Lee, IL Im

この論文をやさしく読む

ひとことで言うと

LLMで国ごとのアンケート回答を模擬するとき、文化差が薄まったり誇張されたりしていないかを測る指標です。

何に役立つ?

合成回答を国際比較に使う前の評価に役立つと考えられます。要旨では従来の分布指標が見逃す国間の差をCDPが捉えたと報告しています。

この研究の面白いところ

各国の回答分布がよく合って見えても、国同士の違いは平板化している場合がありました。

どこまで分かった?

評価は四モデル、三プロンプト法、二調査分野で行われました。CDPには一度の人間回答による較正が必要です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)は、集団の回答分布を推定するため、合成された調査回答者として使われることが増えている。文化をまたぐ調査のシミュレーションでは、各国内の分布を忠実に再現するかだけでなく、国と国の違いが保たれるかも評価すべきである。しかしJensen–Shannonダイバージェンス(JSD)のような既存の距離指標は、国間の違いを直接捉えない。この問題に対処するため、一度だけ人間の回答で較正する、参照情報の少ない診断指標Cultural Divergence Preservation(CDP)を導入する。CDPは国間の差が小さくなることを文化的な平板化、大きくなることを文化の戯画化として捉える。 四つのLLM基盤モデル、三つの人物像に基づくプロンプト法、世界価値観調査(WVS)とビッグファイブ性格検査の二つの調査分野で実験した。結果は、従来の分布忠実度指標とCDPの間に体系的な食い違いがあることを示した。制御実験では、国間の差を弱めたり強めたりするとCDPは単調に変化したが、対応するJSDの変化は比較的小さかった。実際のLLM生成の監査では、DeepPersona-Inspiredというプロンプト法は従来指標に好まれることが多い一方、すべてのモデルと分野の組み合わせで最も強い平板化を示した。CDPは国間の差の縮小や拡大を直接定量化し、従来の忠実度指標を補う。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration. CDP identifies reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature. To evaluate CDP, we conduct experiments across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test. The results reveal a systematic discrepancy between conventional fidelity metrics and CDP. Controlled experiments show that CDP changes monotonically as cross-country divergence is attenuated or amplified, while the corresponding changes in JSD remain relatively small. In our audit of real LLM generations, DeepPersona-Inspired prompting is frequently favored by conventional fidelity metrics but exhibits the strongest flattening in every model--domain block. CDP thus complements fidelity metrics by directly quantifying the attenuation or amplification of cross-country divergence.

著者のコメント

Accepted to the EMNLP 2026 Workshop on Pluralistic AI & NLP (PANDORA)

arXiv ID: 2609.29928 / 要約の誤りについて