人の行動理由まで再現する社会シミュレーションの評価
What People Almost Did: Evaluating LLM Social Simulations Beyond Behavioral Fit
この論文をやさしく読む
ひとことで言うと
人と同じ行動をしただけでは、社会シミュレーションが人の判断理由まで再現したとは言えないと論じています。
何に役立つ?
LLMによる社会シミュレーションを、行動の説明や介入の比較に使う際に、何を確かめる必要があるかを整理するために役立ちます。
この研究の面白いところ
沈黙が無関心からなのか発言抑制からなのかで意味が異なるように、状況・推論・行動の組が対象集団に忠実かという評価軸を提案します。
どこまで分かった?
提案した表現の適切さをどう測るかは未解決問題として提示されています。LLMが示す推論過程が実際の人の理由を正確に表すと実証したという内容ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)に基づく社会シミュレーションは、主に行動の適合度で評価されている。これは、エージェントが、模擬対象の人々の行動や回答分布を再現するかを検証するものである。しかし、シミュレーションに期待される役割は、行動を合わせることにとどまらない。人間の行動を説明し、障壁を診断し、大規模な介入を比較することもできる。こうした用途では、人々が何をしたかだけでなく、なぜそのように行動したのかを理解する必要がある。 したがって、こうした主張を支えるには、行動の適合度だけでは不十分である。行動から、その背後にある推論過程を一意には決められないからである。例えば、沈黙は無関心による場合も、発言を抑えられた結果である場合もある。電話に出ないのは、発信者への不信のためかもしれず、電話を使える機会が限られるためかもしれない。 本論文では、LLMに基づく社会シミュレーションの新たな評価目標として「表現の適切性」を提案する。LLMの推論履歴を活用し、シミュレーションの「状況・推論・行動」の三つ組が、対象とする集団と状況に忠実な形で、行動の背後にある推論過程を保持しているかを測るものである。表現の適切性を解釈可能性やアラインメントの指標と区別し、シミュレーション研究に組み込む方法を提案するとともに、その測定を未解決問題として提示する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
LLM-based social simulations are primarily evaluated for behavioral fit, testing whether agents reproduce the actions or response distributions of the people they are simulating. However, the promise of simulation extends beyond behavioral fit. Simulations can explain human behavior, diagnose barriers, and compare large-scale interventions. These use cases depend on understanding \textit{why} people acted a certain way, not just \textit{what} they did. As a result, behavioral fit is insufficient for these types of claims because behavior underdetermines the reasoning process behind it. For instance, the behavior of staying silent may be due to disinterest or suppressed speech, and not answering a call may be due to distrust of the caller or limited phone access. In this paper, we propose \textit{representational adequacy} as a new evaluation target for LLM-based social simulations. By leveraging LLM reasoning traces, representational adequacy measures whether a simulation's scenario--reasoning--action triples preserve the reasoning process behind the behavior in a way that is faithful to the population and scenarios being simulated. We distinguish representational adequacy from interpretability and alignment metrics, propose ways to integrate it into simulation research, and pose its measurement as an open problem.
arXiv ID: 2609.20055 / 要約の誤りについて