交渉AIは相手の心と会話記録のどちらを追うのか
Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation
この論文をやさしく読む
ひとことで言うと
交渉が自然に進んでも、AIが相手の本当の優先順位を理解しているとは限らないことを、正解の分かる交渉課題で調べました。
何に役立つ?
社会シミュレーションや交渉エージェントを評価するとき、合意率だけでなく相手理解と合意の質を測るために役立ちます。
この研究の面白いところ
相手の発言を変えず自分の利得表だけを変える試験で、相手についての推定が動くことを示し、会話追跡と心的状態の推定を切り分けています。
どこまで分かった?
二つのモデル系列と構成された40の交渉課題の結果です。パレート最適な合意の割合、合意率、相手の信念を予測する精度は異なる指標であり、同一視できません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
言語モデルのエージェントは社会的なやり取りのシミュレーションにますます使われ、その会話記録は、エージェント同士が互いを理解しているかのように読める。本研究では、その見かけが相手の心のモデルに基づくのか、相手が何を言ったかという表面的な記録に基づくのかを問う。両方の問いに厳密な答えを与えられる社会シミュレーションを構築する。対象は40の複数争点交渉で、隠れた選好の重みと完全なパレートフロンティアは構成上既知である。 二つのモデル系列が160の二者組で交渉する。測定前にすべての会話記録を固定し、その後2,880の反実仮想プローブを実施する。証拠をバイト単位で同一に保ちながら、読む側自身の利害、相手の口調、アイデンティティのラベル、再帰の次数という要因を一つずつ動かす。エージェントは社会的には流暢だが、経済的には低調である。二者組の96.2%で合意し、手順違反は0件だった一方、合意のうちパレートフロンティア上にあるのは0.7%にすぎず、利用可能な共同価値の20.5%を取り残し、利害が完全に一致する一つの争点を76.6%の合意で逃した。フロンティアとこの利害一致の争点に関しては、双方が受け入れる集合から無作為に選んだ組み合わせも同程度の成績だった。 プローブは失敗の所在を明らかにする。相手の発言と提案を同じにしたまま、読む側自身の利得表だけを交換すると、推定された最優先事項が15.0パーセントポイント変わった。これは推論よりも自己中心的な投影である。口調の書き換えでは5.3ポイント、アイデンティティのラベルでは0.0ポイント変わった。最も示唆的なのは、エージェントが、相手が自分について何を信じているかを72.5%の割合で予測する一方、その相手の信念自体が正しいのは51.2%にとどまることである。エージェントは、会話の背後にある心よりも、会話をはるかによく追跡する。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-21 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-21 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one another. We ask whether that appearance rests on a model of the partner's mind or on the surface record of what the partner said. We build a social simulation in which both questions have exact answers: 40 multi-issue negotiations whose hidden preference weights and whose full Pareto frontier are known by construction. Two model families negotiate across 160 dyads, every transcript is frozen before any measurement, and 2880 counterfactual probes then hold the evidence byte identical while moving one factor at a time: the reader's own stake, the partner's tone, an identity label, and the order of recursion. The agents are socially fluent and economically poor. They reach agreement in 96.2% of dyads with 0 protocol failures, yet only 0.7% of deals land on the Pareto frontier, they leave 20.5% of the available joint value unclaimed, and they miss the one issue on which their interests are perfectly aligned in 76.6% of deals; on the frontier and on that aligned issue, a package drawn at random from the set both sides would accept does as well. The probes locate the failure. Swapping only the reader's own payoff sheet, while the partner's words and offers stay identical, moves the inferred top priority by 15.0 percentage points, which is egocentric projection rather than inference, while a tone rewrite moves it by 5.3 percentage points and an identity label by 0.0. Most tellingly, an agent predicts what its partner believes about it 72.5% of the time while that partner's belief is itself correct only 51.2% of the time: the agents track the conversation far better than they track the mind behind it.
arXiv ID: 2609.24146 / 要約の誤りについて