arXiv論文メモ
新着一覧
cs.AI · 査読状況未確認

人間社会の多様な意見をLLMが一つの回答へ縮める問題

The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance

Rojin Ziaei

この論文をやさしく読む

ひとことで言うと

国ごとの人間の回答をLLMで模擬すると、平均的な正答率が上がっても意見のばらつきが失われる問題を測った研究。

何に役立つ?

調査回答や社会集団の模擬にLLMを使う際、平均の一致だけでなく意見分布の再現を評価する必要性を示す。

この研究の面白いところ

最も正確なモデルでも人間のばらつきは全体で半分、ナイジェリアでは11%しか残らなかった。

どこまで分かった?

結果はWVSの12か国・10,000件の回答者と質問の組、評価したモデル群に基づく。すべての国や社会調査への同じ程度の収縮を示すものではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデル(LLM)で多様な人々を模擬することは計算社会科学に役立つ可能性があるが、多くの評価は現実の集団内での意見の広がりではなく平均的な回答を採点する。本研究は、六大陸12か国にまたがる世界価値観調査(WVS)の回答者と質問の組10,000件を用い、点ごとの正確さと、予測の標準偏差を人間の標準偏差で割った分散保持率drを同時に測る診断枠組みを作った。ゼロショットの言語モデル11種と、WVSデータで教師あり微調整(SFT)、直接選好最適化(DPO)、群相対方策最適化(GRPO)を用いて微調整した五つの変種を評価した。その結果、アラインメント学習で出力が集団ごとの一つの固定観念に近づく「合意への収縮」という失敗形態を特定した。Llama 3.1 70Bの基盤モデルからTulu 3の各段階までを見ると、最初の教師あり指示調整だけで意見の広がりが半分になり、正確さの改善は小さかった(drは1.22から0.59、正確さは0.9ポイント増)。後段の学習でも広がりは回復しなかった。西洋・高学歴・工業化・富裕・民主主義の国々(WEIRD)とそれ以外の国々との差も生じ、点ごとの正確さを追う調査データでの微調整により深まった。最も正確なモデルであるWVSで微調整したTulu 3 70B-DPOの正確さは57.9%だったが、人間の意見の広がりは全体で半分(dr=0.50)、ナイジェリアでは11%しか保たず、WEIRD諸国では0.70~0.87だった。サンプリング温度を1.0に上げても、微調整した二つのDPOモデルで人間の分布までのWasserstein-1距離は変わらなかった。Qwen 3.5 9BでのGRPOも、正確さを報酬とする場合、分布に合わせた報酬を使う場合のどちらでも広がりを回復しなかった。調整済みモデルと未調整の事前分布を混ぜると、未使用データでdrは0.51から0.62に上がったが、ナイジェリアでは0.36にとどまった。したがって点ごとの正確さだけでは社会の模擬モデルを誤って評価し、現在の追加学習は多様性を合意と引き換えにしている。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-22(UTC)
最新改訂
2026-09-22 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Using large language models (LLMs) to simulate diverse human populations has the potential to transform many aspects of computational social science, yet many evaluations score the average response rather than the spread of opinion within real groups. Here, we develop a diagnostic framework that measures point accuracy alongside dispersion retention, the ratio of predicted to human standard deviation ($\dr$), on 10{,}000 respondent--question pairs from the World Values Survey (WVS) spanning twelve countries and six continents. We evaluate eleven zero-shot language models and five variants fine-tuned on WVS data with SFT, DPO, and GRPO. We identify a failure mode we term \textit{consensus collapse}, where alignment training compresses outputs toward one stereotype per group. Along the post-training trajectory from the Llama~3.1 70B base to the Tulu~3 checkpoints, the first stage, supervised instruction tuning, removes half of the spread with minimal accuracy gain ($\dr$ 1.22 to 0.59; accuracy $+0.9$ points), the later stages do not restore it, and a gap opens between WEIRD and non-WEIRD countries that survey fine-tuning then deepens while pursuing higher point accuracy. The most accurate model (Tulu~3 70B-DPO fine-tuned on WVS, 57.9\%) keeps half the human spread overall ($\dr = 0.50$) and 11\% of it for Nigeria, against 0.70--0.87 for WEIRD countries. Raising the sampling temperature to 1.0 leaves the Wasserstein-1 distance ($\wone$) to human distributions unchanged for both fine-tuned DPO models, and GRPO on Qwen~3.5 9B does not restore the spread under either an accuracy reward or a distribution-shaped reward. Mixing the aligned model with an unaligned prior raises $\dr$ from 0.51 to 0.62 on a held-out split but leaves Nigeria at 0.36. Point accuracy alone therefore misjudges these simulators, and current post-training trades diversity for consensus.

arXiv ID: 2609.25760 / 要約の誤りについて