利用者から圧力を受けた言語モデル60種の振る舞い
Conduct Under Pressure: What Sixty Language Models Do When a User Pushes
この論文をやさしく読む
ひとことで言うと
利用者が強く要求したとき、言語モデル60種が正しい判断を保つか、どう応答するかを同じ場面で比較しました。
何に役立つ?
対話モデルの評価項目を設計し、利用者の圧力に対する応答の違いを点検する材料になります。
この研究の面白いところ
判断を保つ頻度はモデルの世代と、保ち方の様式は提供企業と関連するという、異なる二つの傾向を分けて報告しています。
どこまで分かった?
全モデルに固定した場面を与えた比較です。分類には読み手間の一致に限界があり、実際の利用者との自由な対話全体を代表するとは限りません。
v2のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
利用者が言い張る、懇願する、持ち上げる、悲しみを訴えるといった不快な状況で、言語モデルが正しい事実を引っ込めたり、断るべき文書を書いたり、利用者に金銭的損失をもたらす計画を後押ししたりするかを調べた。13社の60モデルに対し、応答に関係なく全モデルで同じ、複数ターンからなる固定の場面を与えた。各会話は、公開の符号化作業を経て固定した分類基準により、立場を保ったか折れたかという経過と、その保ち方・折れ方の様式を付けた。 モデルが立場を保つかどうかは世代、すなわち新しさと関連した。折れる割合と公開の能力指数とのSpearman相関は−0.64で、提供企業の効果は小さかった。一方、どう立場を保つかは提供企業と関連し、様式を表す17分類のうち六つで、分類全体に対する補正後の置換検定のp値が0.001以下だった。信頼性を満たす分類について、四つの企業別の特徴も報告した。 分類のどの部分に人が必要かも検討した。三社の六つのLLM判定器は、三人の人間の判定者より分類基準を一貫して適用し、Krippendorffのαは0.66対0.46だった。基準の例で使われていない会話について、経過の判定で基準の作成者とのκは0.84~0.91、調停後の人間による基準との一致度は0.83だった。対象を伏せた機械判定は分類の種類を再現できたが、別の読み手がどの分類を同じように適用するかまでは分からなかった。専門家以外にも判断可能な振る舞いでは、人の役割は大量のラベル付けよりも、分類の設計と範囲の設定、小規模な基準データへの責任にあると結論付ける。
v2の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-21(UTC)
- 最新改訂
- 2026-09-23 · v2
- 査読・掲載
- 査読状況未確認
更新履歴
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
We study what LLMs do when a user applies pressure in an uncomfortable situation: a user insists, begs, flatters or grieves, and the model gives up a correct fact, writes a document it should refuse, or cheers a plan that will cost the user money. We send frozen multi-turn scenes, identical for every model regardless of the reply, to 60 models from 13 vendors, and label each transcript with a codebook built by open coding and then frozen: a trajectory (the model held its position or folded) and a manner (how it held or folded). Two findings separate. Whether a model holds tracks its generation, meaning how recent it is: fold rate correlates with a public capability index at Spearman -0.64, with little vendor effect. How it holds tracks the vendor: six of the 17 manner codes sort by vendor at permutation p <= 0.001, corrected across the codebook. We report four vendor profiles on the codes that cleared reliability. We also ask which parts of the labeling need a person. Six LLM coders from three vendors apply the codebook more consistently than three human coders do (Krippendorff's alpha 0.66 against 0.46), agree with the codebook's author on trajectory at kappa 0.84 to 0.91 on transcripts the codebook's examples never touched, and match an adjudicated human reference at 0.83. Blind machine readings recover the codebook's categories but cannot tell which of them a second reader would apply the same way. We conclude that for behavior a non-specialist can judge, the human contribution is authoring and bounding the codes and owning a small reference, not producing labels at volume.
著者のコメント
16 pages, 1 figure, 4 tables. v2 adds a preregistered replication on six further scenes. Code, data and labels: https://github.com/tap2k/modelun/studies/conduct
arXiv ID: 2609.25447 / 要約の誤りについて