arXiv論文メモ
新着一覧
cs.CL / cs.AI · 査読状況未確認

心理尺度で言語モデルの応答傾向を比較する手法

Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling

Yu Sha, Junqi Tao, Dixin Zhou, Yansheng Tu, Mingyang Chen, Xiang Fan, Yang Liu, Mengquan Yang, Jie Lin, Jiahui Fu, Hua Zheng, Benwei Zhang, Zhou Kai

この論文をやさしく読む

ひとことで言うと

心理尺度への回答と回答できなかった項目を使い、言語モデルごとの応答傾向を中国語と英語で比較した研究である。

何に役立つ?

言語モデルの導入時に、回答傾向や回答を拒む範囲を体系的に比較するための評価設計に役立つ。人間の性格をそのまま測るものではない。

この研究の面白いところ

九モデル、七尺度、二言語で五回ずつ評価し、採点できない回答自体にもモデル固有の構造があると示した。

どこまで分かった?

特徴は言語やプロンプトの文脈に依存する。要旨は心理尺度への応答傾向を扱い、モデルに人間と同じ心理特性があるとは示さない。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

大規模言語モデルは人間の意思決定やコミュニケーションを仲介する場面が増えているが、その行動上の規則性を体系的に特徴付けることは難しい。本研究は、言語をまたいだ心理測定的なプロファイル作成の枠組みを構築し、七つの心理尺度を用いて九つの言語モデルを評価した。中国語と英語のそれぞれについて、モデルごとに五回繰り返して尺度を実施した。事前に定めた再試行手順の後も回答が確定しない項目は、欠損値NAとして残す。採点できた回答とNA回答を合わせて分析することで、回答の傾向と自己報告形式を適用できる範囲を捉える。モデルには、向社会的・自己調整的な回答が多く、支配的な態度、関与の放棄、有害な意図の支持が少ないという、共通の調整に影響された傾向がある。それでも、モデルごとの構造的なプロファイルが見られた。NA回答は一様には分布せず、出力が該当しないと扱われた場所、回答が拒否された場所、有効な選択肢に対応付けられない場所を示した。使用言語と提供元はプロファイルの構成と回答可能性に関連した。一方、繰り返し実施した結果の再現性は高く、モデルの同一性も推定できた。人間を参照した分析とプロンプトに対する頑健性の分析は、これらの特徴が文脈に依存することも示した。心理測定のプロファイルと回答可能性を合わせて分析する方法は、導入時の行動上の特徴を定量化する枠組みになる。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-19(UTC)
最新改訂
2026-09-19 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric profiling framework and evaluate nine LLMs using seven psychological instruments, with five repeated administrations per model and language in Chinese and English. Items unresolved after a prespecified retry procedure are retained as NA. Joint analysis of scored and NA responses captures response tendencies and boundaries of self-report applicability. LLMs exhibit structured, model-specific profiles despite a shared alignment-shaped pattern of higher prosocial and self-regulatory responses and lower dominance, disengagement and harmful-intent endorsement. NA responses are structured rather than uniformly distributed, indicating where outputs are treated as inapplicable, refused or cannot be mapped to valid response options. Language condition and provider origin are associated with profile configuration and answerability, whereas repeated administrations show high reproducibility and permit recovery of model identity. Human-reference and prompt-robustness analyses further indicate that these signatures are context dependent. Joint analysis of psychometric profiling and answerability offers a framework for quantifying deployment-level behavioural signatures.

著者のコメント

22 pages, 6 figures

arXiv ID: 2609.22934 / 要約の誤りについて