arXiv論文メモ
新着一覧
cs.CL · 掲載先の記載あり

言語モデルの丁寧さ判定と人間の判断のずれ

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

Rong Wang, Kun Sun, and Yadong Guo

この論文をやさしく読む

ひとことで言うと

言語モデルが文章の丁寧さを判定するとき、人間の判断とどう食い違うかを調べます。

何に役立つ?

丁寧さ判定を使うシステムの評価で、全体の一致率だけでは見落とす偏りを確認する材料になります。

この研究の面白いところ

7モデルの間では人間との一致よりモデル同士の一致が強く、「中立」を出し過ぎる傾向がありました。

どこまで分かった?

評価は2種類の英語データセットで行われました。他言語で同じ傾向があるかは要旨に記載されていません。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

標準的なベンチマークでの性能が高くても、大規模言語モデル(LLM)が社会的な言葉の使い方を人間と同じように評価するかは明らかではない。本研究は、人間による連続的な評価点と3分類のラベルという、相補的な注釈形式を持つ2種類の英語データセットを用い、LLMの丁寧さ判定を評価する。評価した7モデルでは、モデル同士の一致度の方が、モデルと人間の一致度より高かった。表現方略ごとの分析から、モデルと人間の判断の一致は明示的な言語的手掛かりと関係し、一部の関係づくりの方略は判断が食い違う事例でより多く見られることが示唆された。 3分類の課題では、モデルが「中立」ラベルを出し過ぎ、「失礼」ラベルを過少に予測する、体系的な中立への圧縮が見られた。この傾向は、診断用の一部データで専門家の合意を基準にした場合にも続いた。これらの結果は、全体の一致率だけでなく、人間の基準を変えたときにどちらの方向に判断がずれるかも調べる、実用的な言語使用の評価が必要であることを示す。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-24(UTC)
最新改訂
2026-09-24 · v1
査読・掲載
掲載先の記載あり

著者による掲載先の記載:EMNLP2026。出版社での独立確認は未実施です。

arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and three-way categorical labels. Across the seven evaluated models, we find that inter-model agreement is stronger than model--human agreement. Strategy-level analyses suggest that model--human alignment is associated with explicit linguistic cues, while some rapport-building strategies occur more frequently in misaligned cases. In the categorical task, model predictions exhibit systematic neutral compression, characterized by the overproduction of Neutral labels and the underprediction of Impolite labels. This pattern persists when expert consensus is used as the reference on a diagnostic subset. Our findings highlight the need for pragmatic evaluations that go beyond aggregate agreement metrics by examining directional patterns of model--human disagreement across different human references.

arXiv ID: 2609.29001 / 要約の誤りについて