AIアシスタントの政治的回答が利用者の属性でどう変わるか
Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity
この論文をやさしく読む
ひとことで言うと
政治的な質問へのAIの回答や拒否が、話題と利用者の政治的属性によってどう変わるかを調べた研究です。
何に役立つ?
AIアシスタントの監査で、平均的な回答だけでなく、利用者属性ごとの回答・拒否・立場を検査する設計に役立ちます。
この研究の面白いところ
同じシステムでも対照話題と政治的な話題で振る舞いが異なり、版が変わると方針も変化しました。拒否を欠測として捨てずに結果へ含めています。
どこまで分かった?
結果は事前登録した五つの話題、6システム、7,500件の会話に基づきます。政治全般や将来の全バージョンに同じ振る舞いが続くとは示していません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデルを使うAIシステムは、何億人もの人々の政治的な質問に答えている。現在の監査は平均的な利用者に何を答えるかを測るが、実際の振る舞いは動的である。本研究は、政治に関する振る舞いを、話題とシステムが知る利用者の情報に応じて、誰に答えるか、何を言うか、そもそも応答に加わるかを決める方針の集合として捉える。この方針を「発話レジーム」と呼び、開発者が回答、利用者への迎合、拒否の間のトレードオフをどう定めるかを表す。それぞれの費用は話題によって異なる。対話への参加と立場という二つの次元から、五つのレジームの類型を導いた。OpenAI、Anthropic、xAI、Google、Mistral、DeepSeekの6システムを、事前登録した7,500件の複数往復の会話実験で調べた。中絶、カタルーニャ独立、気候変動、ナチズム、利害のない対照話題であるピザにパイナップルを載せることの五つを扱い、利用者の政治的属性は無作為に割り当てた。異なる開発者の二つの言語モデル判定器で各回答を採点し、人間の符号化と照合した。拒否も欠測ではなく結果として扱った。対照話題では全システムが利用者に迎合したため、政治的な話題での抑制は方針によるものと解釈した。意見の対立する話題ではシステムごとに異なるレジームが見られた。中絶ではGPTは全利用者と対話し、その立場を映した。Gemmaは全員を拒否し、Claudeは強く保守的な利用者には35%の割合で答え、ほかの利用者にはほとんど答えず、Grokは保守的な利用者だけに迎合した。気候変動やナチズムのように見解が定まった話題では、五つのシステムがどの利用者に対しても立場を変えなかった。システムは利用者の全体的な思想傾向も推測し、迎合がまだ話していない話題にも及び得た。Grokの二つの版を比較すると、現行の監査では見落とすようなレジームの変化があった。発話レジームはAIの整合性研究、分極化、政治知識、民主主義の質に関わる。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-19(UTC)
- 最新改訂
- 2026-09-19 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-19 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user. I call these policies the system's speech regime, which is how a developer settles the tradeoff between answering, accommodating the user, and refusing, each of which carries a cost that varies by topic. I derive a typology of five regimes from two dimensions, engagement and stance. I test six AI systems (OpenAI, Anthropic, xAI, Google, Mistral, DeepSeek) in a preregistered experiment of 7,500 multi-turn conversations that randomly assign the user's political identity across five topics: abortion, Catalan independence, climate change, Nazism, and a zero-stakes control (pineapple on pizza). Two LLM judges from different developers score every answer, validated against human coding, and refusal is treated as an outcome rather than missing data. Every system accommodates the user on the control topic, showing that political restraint is a policy. On contested topics the systems fall into different regimes: on abortion, GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only. On settled topics such as climate change and Nazism, five systems hold firm for every user. The systems also infer the user's overall ideology, so accommodation can spill over to topics not yet discussed. A comparison of two Grok releases shows the regime changing between versions in a way current audits miss. Speech regimes matter for alignment research and for polarization, political knowledge, and the quality of democracy.
著者のコメント
51 pages, 6 figures. Preregistered; registration at doi:10.5281/zenodo.21135155
arXiv ID: 2609.23039 / 要約の誤りについて