arXiv論文メモ
新着一覧
cs.CY / cs.CL · 査読状況未確認

AIの商品推薦は繰り返し回答やAPIと画面でどう変わるか

"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations

Lucas G. Uberti-Bona Marin and Thales Bertaglia and Giovanni Astante and Bram Rijsbosch and Gijs van Dijck and Anikó Hannák and Gerasimos Spanakis and Konrad Kollnig

この論文をやさしく読む

ひとことで言うと

AIの商品推薦を、回答の繰り返し、利用画面、API、表示される情報源という違いから比較した監査研究です。

何に役立つ?

商品推薦AIを調査する際、一度の回答やAPIだけで消費者の体験を判断しないための調査設計に役立ちます。

この研究の面白いところ

商品への一人称の好みだけでなく、同じ質問で示す情報源のドメインがサービス間でもAPIと画面の間でも大きく異なることを数値化しています。

どこまで分かった?

これらは対象となった回答と条件での観察です。広告が個々の推薦を変えたという因果効果や、推薦商品の品質の優劣を示す結果ではありません。API観察が画面を代表しない点も研究自身の注意事項です。

v1のアブストラクトに基づくAI解説。日本語訳とは別に、用途の解釈を含みます。

アブストラクトの日本語訳

消費者は、何を買うかの助言を得るためにAIチャットボットを使うようになっている。OpenAIやGoogleのような企業が広告を通じてAIを収益化するなかで、こうした助言の偏りや公平性について難しい問題が生じる。そこで本研究では、実際の購買相談を使い、広く使われるチャットボットを監査する。まず、実際の購買相談2,528件からなるデータセットConsumerQを整備する。次に、ChatGPT(チャットボットとAPI)、Google Gemini(チャットボットとAPI)、Google検索(AI Overviews)による商品質問への1,536件の回答を評価する。 商品を推薦する回答のうち、ChatGPTは79%で一人称による商品の好みを表現するのに対し、Geminiは7%、AI Overviewsは2%であることが分かった。また、繰り返し質問すると、推薦商品はしばしば変わる。表示される情報源も大きく異なる。同じ質問に対し、ChatGPTとGeminiの画面が共有するドメインは平均5.4%にとどまり、比較の76.7%では共通ドメインが一つもない。APIが示す情報は対応する画面とも異なり、平均ドメイン重複率はChatGPTで12.0%、Geminiで14.8%であり、公開される情報源情報の種類や層にも違いがある。 これらの知見は、単独の回答もAPIからの観察も、消費者が接する購買助言を代表すると仮定できないことを示す。そのため、AIを介した購買助言の独立した監査では、繰り返しの回答、消費者向けの利用条件、観察対象となる情報源の層を考慮すべきである。

v1の要旨から自動生成。本文の精読・人による確認は未実施。

初稿
2026-09-16(UTC)
最新改訂
2026-09-16 · v1
査読・掲載
査読状況未確認
arXivで読むPDF

更新履歴

取得できた版を表示。版の更新は査読済みを意味しません。過去版の本文差分は未解析です。

原文の要旨

Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this raises difficult questions about the bias and impartiality of such advice. In response, we conduct an AI audit of popular chatbots using real commercial-advice queries. First, we curate a dataset of 2,528 real commercial-advice queries (ConsumerQ). Then, we evaluate 1,536 responses to product queries from popular AI chatbots: ChatGPT (chatbot and API), Google Gemini (chatbot and API), and Google Search (AI Overviews). We find that ChatGPT expresses a first-person product preference in 79% of product-recommending responses, compared with 7% for Gemini and 2% for AI Overviews, while the products recommended often change across repeated requests. Displayed sources vary strongly: for the same query, the ChatGPT and Gemini interfaces share only 5.4% of domains on average, with no domain in common in 76.7% of comparisons. APIs provide a different view from their corresponding interfaces, with mean domain overlaps of 12.0% for ChatGPT and 14.8% for Gemini, and also differ in the types and layers of source information they expose. Our findings show that neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter. Independent audits of AI-mediated commercial advice should therefore account for repeated responses, consumer-facing conditions, and the source layer being observed.

arXiv ID: 2609.18729 / 要約の誤りについて