GUIエージェントは推論を増やしても誘導に弱いのか
A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents
この論文をやさしく読む
ひとことで言うと
買い物画面の初期設定や社会的な誘導が、画面を操作するAIエージェントの選択を変えるか調べています。
何に役立つ?
AIへ意思決定を任せる組織が、画面設計による影響を評価する材料になります。推論を増やせば一律に頑健になるとは限らないことを示します。
この研究の面白いところ
六モデル、3,600エージェント、21,600シミュレーションの無作為化実験です。推論の強化はデフォルトへの影響を減らす一方、社会的影響を使う誘導への感受性を高めました。
どこまで分かった?
オンライン買い物の実験設定における傾向です。モデル規模による違いは探索的分析として報告されており、全てのGUI作業や誘導方式への一般化は示されていません。
v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。
アブストラクトの日本語訳
大規模言語モデル(LLM)に基づくGUIエージェントは、人間の利用者を念頭に設計されたデジタル環境で、利用者に代わって行動する機会が増えている。こうしたグラフィカルユーザーインターフェースは、利用者の行動と意思決定を支援すると同時に、意図的に方向付けるよう設計されてきた。LLMのテキスト出力に現れる行動上の偏りはよく記録されているが、モデルがインターフェースを知覚し、意思決定を実行するエージェントとして動くときに、そのような影響がどう働くかは十分に分かっていない。特に、エージェントに組み込まれつつある推論能力が、この影響への頑健性を高めるかどうかは不明である。 二重過程理論に基づき、LLMベースのGUIエージェントが、自動的な反応を促すタイプ1と熟慮を促すタイプ2のデジタルナッジに影響されるか、また推論設定がその受けやすさをどう変えるかを実証的に調べる。3社の最先端モデル6種類を用い、3,600体のエージェント、計21,600回のシミュレーションによる無作為化オンラインショッピング実験を行ったところ、エージェントは両方のナッジに弱かった。 推論設定は、この二つの効果を逆方向に調整した。自動的なデフォルト選択のナッジへの影響されやすさを低下させる一方、熟慮を介する社会的影響のナッジへの影響されやすさを高めた。したがって、推論を大幅に増やしても頑健性は高まらず、選択肢の提示設計が効果を及ぼす経路が変化した。探索的分析では、この経路の変化がモデル規模によって系統的に構造化されることも示された。本研究は、ナッジへの感受性をエージェント型AIの行動特性として示すとともに、自律エージェントに意思決定を委ねる組織にとって、インターフェース設計がガバナンス上の課題となることを示す。
v1の要旨から自動生成。本文の精読・人による確認は未実施。
- 初稿
- 2026-09-17(UTC)
- 最新改訂
- 2026-09-17 · v1
- 査読・掲載
- 査読状況未確認
更新履歴
- v1 2026-09-17 この版を読む
取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。
原文の要旨
LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in the textual outputs of LLMs are well-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions---and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it. Drawing on Dual-Process Theory, we empirically investigate whether LLM-based GUI agents are susceptible to automatic (Type 1) and reflective (Type 2) digital nudges, and how their reasoning configuration moderates this susceptibility. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect. Exploratory analysis further showed this redirection to be systematically structured by model scale. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents.
著者のコメント
Preprint of a manuscript completed in November, 2025
arXiv ID: 2609.19843 / 要約の誤りについて