人格攻撃への応答から大規模言語モデルの議論行動を評価
Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks
この論文をやさしく読む
ひとことで言うと
政治討論で人格攻撃を受けたとき、人間と大規模言語モデルの応答戦略がどう違うかを比較した研究。
何に役立つ?
議論を行うAIを、論理的な正しさに加えて実際の対話での戦略の幅から評価する際の参考になる。
この研究の面白いところ
人間の政治討論にある防御戦略とモデルの生成対話を比べ、多くのモデルが論理的な防御に偏ることを報告した。
どこまで分かった?
比較は政治対話と米国大統領選討論のコーパスに基づく。安全性の微調整による制限は著者らの解釈として述べられている。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルは説得的な対話を行う議論エージェントとして使われるようになっており、人間の対話相手と比べた議論能力を厳密に評価する必要がある。本研究は、従来は誤謬として退けられがちな人格攻撃、すなわち人身攻撃に注目する。政治的な説得対話では、話者の信頼性や人格が命題の内容に匹敵するほど重要な役割を果たすためである。具体的には、現代の大規模言語モデルが、そのような攻撃を戦略的に用い、応答する人間の能力を再現できるかを調べる。自然言語による政治対話のコーパスを分析し、話者の信頼性が中心となる議論で人間が自然に用いる防御戦略を見いだして、対話ゲームに整理した。実証評価では、米国大統領選の討論を収めたElecDeb60to16-fallacyコーパスを基準に、モデルが生成した対話と人間の討論者の防御戦略の種類を比較した。結果として大きな違いが見られた。ほとんどの大規模言語モデルは論理的な防御を硬直的に優先し、政治的な議論で有効な応答になり得る話者の信頼性に関わる反撃を活用できなかった。著者らは、現在の安全性のための微調整がモデルの戦略的な行動の範囲を制限し、人格をめぐる争いが単なる誤謬ではなく通常のやり取りとして期待される分野では、自然な対話に十分参加できなくしていると論じる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-23(UTC)
- 最新改訂
- 2026-09-23 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-23 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game. Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse. We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.
著者のコメント
Accepted to COMMA 2026
arXiv ID: 2609.28673 / 要約の誤りについて