言語モデル集団は人間の議論の合意率を過大評価する
Language-model groups overstate consensus when replaying human deliberation on a reasoning task
この論文をやさしく読む
ひとことで言うと
人間の討論を言語モデルの集団で再現すると、合意が実際より多く成立してしまうことを調べます。
何に役立つ?
人間の集団意思決定をAIで模擬する研究で、合意率の推定を評価する材料です。100の人間集団を、討論前の回答に対応するエージェント集団と比較します。
この研究の面白いところ
不参加者の扱いなど測定の違いを調整しても、合意率の差が約34〜44ポイント残りました。推論モードでは、誤った答えにほぼ全員が同意する場合もありました。
どこまで分かった?
Wason推論課題という特定の設定の結果です。合意と正答は別であり、要旨はあらゆる討論で同じ差が出ると結論していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
全員一致の割合は集団的認知の指標として扱われることが多いが、その値は参加や最終状態をどのように操作的に定義するかに左右される。本研究では、評価用に留保した人間のウェイソン課題の100グループを、対応する大規模言語モデル(LLM)のエージェント集団で再現した。各参加者の議論前の回答を初期信念とするエージェントを1体ずつ用意し、エージェントと人間を同じコードで採点した。人間の採点定義を変えると合意率の推定値は24.0%から57.0%に広がった。参加者のおよそ5人に1人は一度も投稿しなかったのに対し、エージェントはほぼ常に投稿した。 盲検解除後に行った2つの感度分析でも、エージェント集団の合意率は人間より高かった。回答提出に基づく比較(n=98)では、チャットモードと推論モードの差はそれぞれ34.0、43.9パーセントポイントであり、参加状況を一致させた比較(n=45)では34.1、44.4ポイントだった。この2つの相補的な方法は、異なる測定上の非対称性を減らしながら、差の推定値が0.5パーセントポイント以内で一致した。 早期停止を行わなくても、また記憶可能な正解を除くよう課題を再パラメータ化しても、この差は残った。後者では推論モードの集団はほぼ全員一致したが、その大半は誤答での一致だった。シミュレーション上の合意は集団の正答率と連動せず、この設定では、初期信念を固定したエージェント集団は人間集団の結果分布の偏った推定量となった。これらの分析は、人間の熟議結果を模擬集団で推定する際の評価に、採点方法を明示した基盤を与える。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.
著者のコメント
37 pages, 4 figures. Preregistration: https://osf.io/5jp7s . Code and data: https://doi.org/10.5281/zenodo.21318346
arXiv ID: 2609.20543 / 要約の誤りについて