言語モデルの道徳観回答は指示でどこまで変わるか
Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30
この論文をやさしく読む
ひとことで言うと
AIに人間向けの道徳観アンケートを答えさせたとき、回答が何を反映するのかを調べています。人物像を指定するだけで、人間集団への近さだけでなく、質問内容への反応の仕方まで変わりました。
何に役立つ?
言語モデルの価値観を質問紙で測る研究を評価する際に役立ちます。人間に近い平均値だけでなく、各質問の内容に応じた回答になっているかも確認する必要があると分かります。
この研究の面白いところ
人間の回答分布を教えない北欧の人物設定でも、回答の距離が縮まりました。一方、内部活性への介入は狙った基盤を調整するより、回答全体を平坦化したという違いがあります。
どこまで分かった?
対象は6モデルとノルウェー語MFQ-30で、ActAddも固定中間層・1組の対という設定です。44~77%は距離の二乗の改善であり、道徳性や人間理解そのものが改善した割合ではありません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
近年、人間向けの心理測定質問紙を大規模言語モデルに適用し、道徳や価値観のプロファイルを引き出す研究が行われている。しかし、こうした測定器具がモデル内の何らかの安定したものを測っているのか、また得られるプロファイルを目標とする人間集団に近づけられるのかは明らかでない。本研究では、重みが公開された6つの大規模言語モデルにノルウェー語版道徳基盤質問紙(MFQ-30)を実施し、その基盤別プロファイルを、ノルウェーの回答者1,282人の標本と比較する。プロンプトによるペルソナ誘導と、活性値レベルのActAddという2つの介入を検証する。 使用した注意確認課題で、モデルの半数は質問紙の内容に即して応答する。残りの半数は、一様な回答や中央の選択肢に偏る回答を初期状態で出し、平均では人間に近く見えるものの、項目内容を反映していない。人間標本の分布情報を一切用いずに作成した、中立的な北欧の回答者というペルソナは、内容に即して応答するモデルを、マハラノビス距離の二乗d²で44~77%、ノルウェー人の平均に近づける。固定した中間層で1組の対を用いるActAddは、個別の道徳基盤を誘導するのではなく、基盤別のプロファイルを平坦化する。少なくとも1つのモデルでは、プロファイルを変化させるのと同じペルソナが、ベースラインでは見られなかった内容に即した応答も生じさせる。これは、Peereboomら(2025年)が警告する「認知的幻影」の具体例である。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-18(UTC)
- 最新改訂
- 2026-09-18 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-18 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis $d^2$. One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.
著者のコメント
13 pages, 4 figures, 7 tables. Awarded best Paper Award at WNNLP 2026 (University of Oslo). Proceedings: https://www.uio.no/studier/emner/matnat/ifi/IN5550/v26/final-exam/wnnlp2026_proceedings.pdf
arXiv ID: 2609.21636 / 要約の誤りについて