arXiv論文メモ
新着一覧
cs.CL · 査読状況未確認

合成人口の調査回答が研究に使えるかを検証する

Artificial Societies Benchmark: A Validation Framework for Synthetic Research

Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis, James K. He

この論文をやさしく読む

ひとことで言うと

AIが作った調査回答を研究に使えるか、平均だけでなく個人差や回答間の関係まで評価する枠組みを作った。

何に役立つ?

合成調査を利用する研究で、分析目的に必要な妥当性を確認し、人間のデータを追加すべき箇所を探せる。

この研究の面白いところ

回答者の情報を増やすことが、モデルによって改善にも悪化にもつながった。

どこまで分かった?

比較は20の人間の情報源と9モデルに基づく。ある検査で良い結果でも、別の妥当性が保証されるわけではない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

合成調査は平均的な回答を再現しても、人々の違い、回答同士の関係、条件を変えたときの反応を誤って表すことがある。本研究は、合成人口が予定する分析に使えるかを研究者が評価するためのArtificial Societies Benchmarkを提案する。枠組みは、内部妥当性、構成概念妥当性、外的妥当性にまたがる11の検査を組み合わせ、人間の回答を含む20の情報源を使い、9つの言語モデルを比較する。研究での用途ごとに必要な証拠を対応付け、回答者について与える情報によって結果がどう変わるかも調べる。 一つの領域で高成績でも、ほかの領域で忠実とは限らない。モデルはしばしば回答が一貫しすぎ、回答尺度を狭く使い、特性間の関係を変えてしまった。回答者の詳しいプロフィールは、モデルによって予測を改善する場合も悪化させる場合もあった。得られた評価表は、合成人口のどの面が分析に使えるか、どこに追加の人間のデータが必要かを研究者が見極める助けになる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.

著者のコメント

36 pages, 9 figures, 9 tables

arXiv ID: 2609.30030 / 要約の誤りについて